LungCURE: Benchmarking Multimodal Real-World Clinical Reasoning for Precision Lung Cancer Diagnosis and Treatment
AI 摘要
LungCURE提出,针对肺癌诊疗,构建多模态临床推理基准,并提出LCAgent框架。
主要贡献
- 构建 LungCURE 多模态肺癌诊疗基准
- 提出 LCAgent 多智能体框架,提升诊疗决策
- 验证 LLMs 在复杂医学推理上的能力差异
方法论
构建包含真实临床数据的多模态基准,利用多智能体框架LCAgent辅助LLM进行符合指南的肺癌诊疗。
原文摘要
Lung cancer clinical decision support demands precise reasoning across complex, multi-stage oncological workflows. Existing multimodal large language models (MLLMs) fail to handle guideline-constrained staging and treatment reasoning. We formalize three oncological precision treatment (OPT) tasks for lung cancer, spanning TNM staging, treatment recommendation, and end-to-end clinical decision support. We introduce LungCURE, the first standardized multimodal benchmark built from 1,000 real-world, clinician-labeled cases across more than 10 hospitals. We further propose LCAgent, a multi-agent framework that ensures guideline-compliant lung cancer clinical decision-making by suppressing cascading reasoning errors across the clinical pathway. Experiments reveal large differences across various large language models (LLMs) in their capabilities for complex medical reasoning, when given precise treatment requirements. We further verify that LCAgent, as a simple yet effective plugin, enhances the reasoning performance of LLMs in real-world medical scenarios.