A Systematic Study of Retrieval Pipeline Design for Retrieval-Augmented Medical Question Answering
AI 摘要
系统性评估检索增强医学问答,分析不同检索策略对性能的影响。
主要贡献
- 系统性评估了检索增强医学问答中不同检索组件的影响。
- 证明了检索增强显著提升零样本医学问答性能。
- 分析了检索效果和计算成本之间的权衡。
方法论
使用MedQA USMLE基准和教材知识库,评估不同语言模型、嵌入模型、检索策略、查询重构和交叉编码器重排序的组合。
原文摘要
Large language models (LLMs) have demonstrated strong capabilities in medical question answering; however, purely parametric models often suffer from knowledge gaps and limited factual grounding. Retrieval-augmented generation (RAG) addresses this limitation by integrating external knowledge retrieval into the reasoning process. Despite increasing interest in RAG-based medical systems, the impact of individual retrieval components on performance remains insufficiently understood. This study presents a systematic evaluation of retrieval-augmented medical question answering using the MedQA USMLE benchmark and a structured textbook-based knowledge corpus. We analyze the interaction between language models, embedding models, retrieval strategies, query reformulation, and cross-encoder reranking within a unified experimental framework comprising forty configurations. Results show that retrieval augmentation significantly improves zero-shot medical question answering performance. The best-performing configuration was dense retrieval with query reformulation and reranking achieved 60.49% accuracy. Domain-specialized language models were also found to better utilize retrieved medical evidence than general-purpose models. The analysis further reveals a clear tradeoff between retrieval effectiveness and computational cost, with simpler dense retrieval configurations providing strong performance while maintaining higher throughput. All experiments were conducted on a single consumer-grade GPU, demonstrating that systematic evaluation of retrieval-augmented medical QA systems can be performed under modest computational resources.