MedRCube: A Multidimensional Framework for Fine-Grained and In-Depth Evaluation of MLLMs in Medical Imaging
AI 摘要
MedRCube提出了一个多维度医学影像MLLM评估框架,揭示了现有评估的局限性,并发现了影响临床信任度的因素。
主要贡献
- 提出了多维度、细粒度的医学影像MLLM评估框架MedRCube
- 揭示了现有评估方法的局限性
- 发现了捷径行为与诊断性能之间的关联,并引发了对临床信任度的担忧
方法论
构建了两阶段的系统性评估流程,通过实例化MedRCube,对33个MLLM进行基准测试,并引入可信度评估子集。
原文摘要
The potential of Multimodal Large Language Models (MLLMs) in domain of medical imaging raise the demands of systematic and rigorous evaluation frameworks that are aligned with the real-world medical imaging practice. Existing practices that report single or coarse-grained metrics are lack the granularity required for specialized clinical support and fail to assess the reliability of reasoning mechanisms. To address this, we propose a paradigm shift toward multidimensional, fine-grained and in-depth evaluation. Based on a two-stage systematic construction pipeline designed for this paradigm, we instantiate it with MedRCube. We benchmark 33 MLLMs, \textit{Lingshu-32B} achieve top-tier performance. Crucially, MedRCube exposes a series of pronounced insights inaccessible under prior evaluation settings. Furthermore, we introduce a credibility evaluation subset to quantify reasoning credibility, uncover a highly significant positive association between shortcut behavior and diagnostic task performance, raising concerns for clinically trustworthy deployment. The resources of this work can be found at https://github.com/F1mc/MedRCube.