LLM Reasoning 相关度: 9/10

Beyond Literal Summarization: Redefining Hallucination for Medical SOAP Note Evaluation

Bhavik Vachhani, Kush Shrisvastava, Pranshu Nema, Sai Chiranthan
arXiv: 2604.14829v1 发布: 2026-04-16 更新: 2026-04-16

AI 摘要

该论文重新定义了医疗SOAP报告评估中的幻觉,提出临床推理感知的评估方法。

主要贡献

  • 指出传统词汇忠实度评估方法在医疗领域存在偏差,高估了幻觉率。
  • 揭示了被标记为幻觉的很多情况实际上是合法的临床转换,如同义词映射和诊断推断。
  • 提出基于临床知识和本体的评估方法,更准确地评估LLM在医疗领域的表现。

方法论

分析传统评估方法的缺陷,提出基于医学本体和校准prompt的推理感知评估方法,并进行实验验证。

原文摘要

Evaluating large language models (LLMs) for clinical documentation tasks such as SOAP note generation remains challenging. Unlike standard summarization, these tasks require clinical abstraction, normalization of colloquial language, and medically grounded inference. However, prevailing evaluation methods including automated metrics and LLM as judge frameworks rely on lexical faithfulness, often labeling any information not explicitly present in the transcript as hallucination. We show that such approaches systematically misclassify clinically valid outputs as errors, inflating hallucination rates and distorting model assessment. Our analysis reveals that many flagged hallucinations correspond to legitimate clinical transformations, including synonym mapping, abstraction of examination findings, diagnostic inference, and guideline consistent care planning. By aligning evaluation criteria with clinical reasoning through calibrated prompting and retrieval grounded in medical ontologies we observe a significant shift in outcomes. Under a lexical evaluation regime, the mean hallucination rate is 35%, heavily penalizing valid reasoning. With inference aware evaluation, this drops to 9%, with remaining cases reflecting genuine safety concerns. These findings suggest that current evaluation practices over penalize valid clinical reasoning and may measure artifacts of evaluation design rather than true errors, underscoring the need for clinically informed evaluation in high context domains like medicine.

标签

LLM 医疗 SOAP报告 幻觉评估 临床推理

arXiv 分类

cs.AI