Multimodal Learning 相关度: 9/10

MedLayBench-V: A Large-Scale Benchmark for Expert-Lay Semantic Alignment in Medical Vision Language Models

Han Jang, Junhyeok Lee, Heeseong Eum, Kyu Sung Choi
arXiv: 2604.05738v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

MedLayBench-V是一个用于医学视觉语言模型专家-患者语义对齐的大规模基准数据集。

主要贡献

  • 构建了首个大规模医学图像专家-患者语义对齐多模态基准数据集MedLayBench-V
  • 提出了基于结构化概念的细化(SCGR)的数据集构建方法,保证语义等价
  • 为训练和评估能够弥合医患沟通鸿沟的Med-VLMs提供了验证基础

方法论

通过SCGR流程,结合UMLS CUI和微观实体约束构建数据集,确保专家术语和通俗表达的语义一致性。

原文摘要

Medical Vision-Language Models (Med-VLMs) have achieved expert-level proficiency in interpreting diagnostic imaging. However, current models are predominantly trained on professional literature, limiting their ability to communicate findings in the lay register required for patient-centered care. While text-centric research has actively developed resources for simplifying medical jargon, there is a critical absence of large-scale multimodal benchmarks designed to facilitate lay-accessible medical image understanding. To bridge this resource gap, we introduce MedLayBench-V, the first large-scale multimodal benchmark dedicated to expert-lay semantic alignment. Unlike naive simplification approaches that risk hallucination, our dataset is constructed via a Structured Concept-Grounded Refinement (SCGR) pipeline. This method enforces strict semantic equivalence by integrating Unified Medical Language System (UMLS) Concept Unique Identifiers (CUIs) with micro-level entity constraints. MedLayBench-V provides a verified foundation for training and evaluating next-generation Med-VLMs capable of bridging the communication divide between clinical experts and patients.

标签

医学视觉语言模型 多模态学习 语义对齐 基准数据集 患者沟通

arXiv 分类

cs.CL