Multimodal Learning 相关度: 9/10

Semantic-Topological Graph Reasoning for Language-Guided Pulmonary Screening

Chenyu Xue, Yiran Liu, Mian Zhou, Jionglong Su, Zhixiang Lu
arXiv: 2604.05620v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

提出STGR框架,结合LLM和视觉基础模型,用于语言引导的肺部筛查,效果显著。

主要贡献

  • 提出Text-to-Vision Intent Distillation (TVID) 模块
  • 将掩码选择问题转化为动态图推理问题
  • 引入Selective Asymmetric Fine-Tuning (SAFT) 策略

方法论

结合LLaMA-3-V和MedSAM,通过TVID提取诊断指导,图推理解决解剖模糊,SAFT策略微调。

原文摘要

Medical image segmentation driven by free-text clinical instructions is a critical frontier in computer-aided diagnosis. However, existing multimodal and foundation models struggle with the semantic ambiguity of clinical reports and fail to disambiguate complex anatomical overlaps in low-contrast scans. Furthermore, fully fine-tuning these massive architectures on limited medical datasets invariably leads to severe overfitting. To address these challenges, we propose a novel Semantic-Topological Graph Reasoning (STGR) framework for language-guided pulmonary screening. Our approach elegantly synergizes the reasoning capabilities of large language models (LLaMA-3-V) with the zero-shot delineation of vision foundation models (MedSAM). Specifically, we introduce a Text-to-Vision Intent Distillation (TVID) module to extract precise diagnostic guidance. To resolve anatomical ambiguity, we formulate mask selection as a dynamic graph reasoning problem, where candidate lesions are modeled as nodes and edges capture spatial and semantic affinities. To ensure deployment feasibility, we introduce a Selective Asymmetric Fine-Tuning (SAFT) strategy that updates less than 1% of the parameters. Rigorous 5-fold cross-validation on the LIDC-IDRI and LNDb datasets demonstrates that our framework establishes a new state-of-the-art. Notably, it achieves an 81.5% Dice Similarity Coefficient (DSC) on LIDC-IDRI, outperforming leading LLM-based tools like LISA by over 5%. Crucially, our SAFT strategy acts as a powerful regularizer, yielding exceptional cross-fold stability (0.6% DSC variance) and paving the way for robust, context-aware clinical deployment.

标签

医学图像分割 LLM 图推理 多模态学习

arXiv 分类

cs.CV cs.AI