QuarkMedSearch: A Long-Horizon Deep Search Agent for Exploring Medical Intelligence
AI 摘要
QuarkMedSearch通过数据构建、训练策略和评估基准提升Agent在中文医疗深度搜索领域的性能。
主要贡献
- 构建了大规模医疗知识图谱和实时在线探索结合的长程医疗深度搜索训练数据
- 提出了两阶段SFT和RL训练策略,提升模型规划、工具调用和反思能力
- 构建了QuarkMedSearch Benchmark,用于评估模型在医疗深度搜索任务中的性能
方法论
结合医疗知识图谱和在线探索构建数据,使用SFT和RL两阶段训练,并用人工验证的基准评估。
原文摘要
As agentic foundation models continue to evolve, how to further improve their performance in vertical domains has become an important challenge. To this end, building upon Tongyi DeepResearch, a powerful agentic foundation model, we focus on the Chinese medical deep search scenario and propose QuarkMedSearch, systematically exploring a full-pipeline approach spanning medical multi-hop data construction, training strategies, and evaluation benchmarks to further push and assess its performance upper bound in vertical domains. Specifically, for data synthesis, to address the scarcity of deep search training data in the medical domain, we combine a large-scale medical knowledge graph with real-time online exploration to construct long-horizon medical deep search training data; for post-training, we adopt a two-stage SFT and RL training strategy that progressively enhances the model's planning, tool invocation, and reflection capabilities required for deep search, while maintaining search efficiency; for evaluation, we collaborate with medical experts to construct the QuarkMedSearch Benchmark through rigorous manual verification. Experimental results demonstrate that QuarkMedSearch achieves state-of-the-art performance among open-source models of comparable scale on the QuarkMedSearch Benchmark, while also maintaining strong competitiveness on general benchmarks.