Multimodal Learning 相关度: 9/10

Lost in the Hype: Revealing and Dissecting the Performance Degradation of Medical Multimodal Large Language Models in Image Classification

Xun Zhu, Fanbin Mo, Xi Chen, Kaili Zheng, Shaoshuai Yang, Yiming Shi, Jian Gao, Miao Li, Ji Wu
arXiv: 2604.08333v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

医学多模态大语言模型在图像分类任务中表现不如传统模型,论文分析了其性能退化的原因。

主要贡献

  • 揭示了医学多模态大语言模型在医学图像分类任务中性能退化的现象
  • 通过特征探测分析了性能退化的四个主要原因
  • 提出了量化指标来评估特征演变的健康程度

方法论

在14个开源医学MLLM和3个数据集上进行实验,采用特征探测技术,逐模块逐层追踪视觉特征的信息流。

原文摘要

The rise of multimodal large language models (MLLMs) has sparked an unprecedented wave of applications in the field of medical imaging analysis. However, as one of the earliest and most fundamental tasks integrated into this paradigm, medical image classification reveals a sobering reality: state-of-the-art medical MLLMs consistently underperform compared to traditional deep learning models, despite their overwhelming advantages in pre-training data and model parameters. This paradox prompts a critical rethinking: where exactly does the performance degradation originate? In this paper, we conduct extensive experiments on 14 open-source medical MLLMs across three representative image classification datasets. Moving beyond superficial performance benchmarking, we employ feature probing to track the information flow of visual features module-by-module and layer-by-layer throughout the entire MLLM pipeline, enabling explicit visualization of where and how classification signals are distorted, diluted, or overridden. As the first attempt to dissect classification performance degradation in medical MLLMs, our findings reveal four failure modes: 1) quality limitation in visual representation, 2) fidelity loss in connector projection, 3) comprehension deficit in LLM reasoning, and 4) misalignment of semantic mapping. Meanwhile, we introduce quantitative scores that characterize the healthiness of feature evolution, enabling principled comparisons across diverse MLLMs and datasets. Furthermore, we provide insightful discussions centered on the critical barriers that prevent current medical MLLMs from fulfilling their promised clinical potential. We hope that our work provokes rethinking within the community-highlighting that the road from high expectations to clinically deployable MLLMs remains long and winding.

标签

Multimodal Learning Medical Imaging Large Language Models

arXiv 分类

cs.CV cs.AI cs.LG