Yale-DM-Lab at ArchEHR-QA 2026: Deterministic Grounding and Multi-Pass Evidence Alignment for EHR Question Answering
AI 摘要
Yale-DM-Lab在ArchEHR-QA 2026上的EHR问答系统,利用模型集成和多步证据对齐提升性能。
主要贡献
- 利用模型集成和投票策略提高性能
- 将完整临床答案段落作为证据对齐的上下文
- 分析了证据对齐的限制主要在于推理能力
方法论
使用Claude Sonnet 4和GPT-4o进行问题重构,使用Azure模型集成进行证据识别、答案生成和对齐,结合few-shot prompting和投票策略。
原文摘要
We describe the Yale-DM-Lab system for the ArchEHR-QA 2026 shared task. The task studies patient-authored questions about hospitalization records and contains four subtasks (ST): clinician-interpreted question reformulation, evidence sentence identification, answer generation, and evidence-answer alignment. ST1 uses a dual-model pipeline with Claude Sonnet 4 and GPT-4o to reformulate patient questions into clinician-interpreted questions. ST2-ST4 rely on Azure-hosted model ensembles (o3, GPT-5.2, GPT-5.1, and DeepSeek-R1) combined with few-shot prompting and voting strategies. Our experiments show three main findings. First, model diversity and ensemble voting consistently improve performance compared to single-model baselines. Second, the full clinician answer paragraph is provided as additional prompt context for evidence alignment. Third, results on the development set show that alignment accuracy is mainly limited by reasoning. The best scores on the development set reach 88.81 micro F1 on ST4, 65.72 macro F1 on ST2, 34.01 on ST3, and 33.05 on ST1.