Multimodal Learning 相关度: 9/10

Enhancing Reinforcement Learning for Radiology Report Generation with Evidence-aware Rewards and Self-correcting Preference Learning

Qin Zhou, Guoyan Liang, Qianyi Yang, Jingyuan Chen, Sai Wu, Chang Yao, Zhe Wang
arXiv: 2604.13598v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

提出ESC-RL,通过证据感知奖励和自校正偏好学习增强放射报告生成。

主要贡献

  • 提出了Group-wise Evidence-aware Alignment Reward (GEAR)
  • 提出了Self-correcting Preference Learning (SPL) 策略
  • 在胸部X射线数据集上取得了state-of-the-art的性能

方法论

利用GEAR提供证据感知反馈,SPL自动构建偏好数据集并使用LLM合成报告,实现临床真实性和自我改进。

原文摘要

Recent reinforcement learning (RL) approaches have advanced radiology report generation (RRG), yet two core limitations persist: (1) report-level rewards offer limited evidence-grounded guidance for clinical faithfulness; and (2) current methods lack an explicit self-improving mechanism to align with clinical preference. We introduce clinically aligned Evidence-aware Self-Correcting Reinforcement Learning (ESC-RL), comprising two key components. First, a Group-wise Evidence-aware Alignment Reward (GEAR) delivers group-wise, evidence-aware feedback. GEAR reinforces consistent grounding for true positives, recovers missed findings for false negatives, and suppresses unsupported content for false positives. Second, a Self-correcting Preference Learning (SPL) strategy automatically constructs a reliable, disease-aware preference dataset from multiple noisy observations and leverages an LLM to synthesize refined reports without human supervision. ESC-RL promotes clinically faithful, disease-aligned reward and supports continual self-improvement during training. Extensive experiments on two public chest X-ray datasets demonstrate consistent gains and state-of-the-art performance.

标签

强化学习 放射报告生成 自然语言处理 医学影像

arXiv 分类

cs.LG stat.ME