Multimodal Learning 相关度: 9/10

MApLe: Multi-instance Alignment of Diagnostic Reports and Large Medical Images

Felicia Bader, Philipp Seeböck, Anastasia Bartashova, Ulrike Attenberger, Georg Langs
arXiv: 2604.13970v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

MApLe提出了一种多任务多实例视觉语言对齐方法,解决医学影像报告中细微病理发现与图像关联的难题。

主要贡献

  • 提出了MApLe模型,用于诊断报告和医学图像的多实例对齐
  • 解耦了解剖区域和诊断发现的概念
  • 基于patch的方式将局部图像信息与句子联系起来

方法论

MApLe模型包含文本嵌入、基于解剖结构的图像编码器和多实例对齐,以实现图像区域和诊断发现的对齐。

原文摘要

In diagnostic reports, experts encode complex imaging data into clinically actionable information. They describe subtle pathological findings that are meaningful in their anatomical context. Reports follow relatively consistent structures, expressing diagnostic information with few words that are often associated with tiny but consequential image observations. Standard vision language models struggle to identify the associations between these informative text components and small locations in the images. Here, we propose "MApLe", a multi-task, multi-instance vision language alignment approach that overcomes these limitations. It disentangles the concepts of anatomical region and diagnostic finding, and links local image information to sentences in a patch-wise approach. Our method consists of a text embedding trained to capture anatomical and diagnostic concepts in sentences, a patch-wise image encoder conditioned on anatomical structures, and a multi-instance alignment of these representations. We demonstrate that MApLe can successfully align different image regions and multiple diagnostic findings in free-text reports. We show that our model improves the alignment performance compared to state-of-the-art baseline models when evaluated on several downstream tasks. The code is available at https://github.com/cirmuw/MApLe.

标签

多模态学习 视觉语言模型 医学影像 报告生成

arXiv 分类

cs.CV