Multimodal Learning 相关度: 9/10

Dialectic-Med: Mitigating Diagnostic Hallucinations via Counterfactual Adversarial Multi-Agent Debate

Zhixiang Lu, Jionglong Su
arXiv: 2604.11258v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

Dialectic-Med通过多智能体对抗辩论,显著缓解医疗多模态大模型诊断中的幻觉问题。

主要贡献

  • 提出 Dialectic-Med 多智能体辩论框架,解决医疗领域 MLLM 的幻觉问题
  • 引入视觉证伪模块,主动检索矛盾视觉证据挑战诊断假设
  • 通过加权共识图解决智能体冲突,提高诊断可靠性

方法论

构建Proponent, Opponent, Mediator三个角色智能体,通过对抗辩论和视觉证据校验,提升诊断准确性和可信度。

原文摘要

Multimodal Large Language Models (MLLMs) in healthcare suffer from severe confirmation bias, often hallucinating visual details to support initial, potentially erroneous diagnostic hypotheses. Existing Chain-of-Thought (CoT) approaches lack intrinsic correction mechanisms, rendering them vulnerable to error propagation. To bridge this gap, we propose Dialectic-Med, a multi-agent framework that enforces diagnostic rigor through adversarial dialectics. Unlike static consensus models, Dialectic-Med orchestrates a dynamic interplay between three role-specialized agents: a proponent that formulates diagnostic hypotheses; an opponent equipped with a novel visual falsification module that actively retrieves contradictory visual evidence to challenge the Proponent; and a mediator that resolves conflicts via a weighted consensus graph. By explicitly modeling the cognitive process of falsification, our framework guarantees that diagnostic reasoning is tightly grounded in verified visual regions. Empirical evaluations on MIMIC-CXR-VQA, VQA-RAD, and PathVQA demonstrate that Dialectic-Med not only achieves state-of-the-art performance but also fundamentally enhances the trustworthiness of the reasoning process. Beyond accuracy, our approach significantly enhances explanation faithfulness and decisively mitigates hallucinations, establishing a new standard over single-agent baselines.

标签

医疗诊断 多模态大模型 智能体 对抗辩论 幻觉缓解

arXiv 分类

cs.CL