Beyond Stochastic Exploration: What Makes Training Data Valuable for Agentic Search
AI 摘要
提出Hierarchical Experience (HiExp)框架,提升LLM搜索Agent的训练效率和稳定性。
主要贡献
- 提出了HiExp框架,利用层级经验知识正则化探索
- 通过对比分析和多层聚类提取经验知识
- 验证了HiExp在agentic search和数学推理上的有效性和泛化性
方法论
通过对比分析和多层聚类将原始推理轨迹转化为层级经验知识,利用经验对齐训练正则化随机探索。
原文摘要
Reinforcement learning (RL) has become an effective approach for advancing the reasoning capabilities of large language models (LLMs) through the strategic integration of external search engines. However, current RL-based search agents often rely on a process of stochastic exploration guided by carefully crafted outcome rewards, leading to inefficient reasoning trajectories and unstable training. To address these issues, we propose a novel framework, Hierarchical Experience (HiExp), to enhance the performance and training stability of search agents. Specifically, we extract empirical knowledge through contrastive analysis and a multi-level clustering mechanism, transforming raw reasoning trajectories into hierarchical experience knowledge. By leveraging experience-aligned training, we effectively regularize stochastic exploration, evolving it into a strategic and experience-driven search process. Extensive evaluations on multiple complex agentic search and mathematical reasoning benchmarks demonstrate that our approach not only achieves substantial performance gains but also exhibits strong cross-task and cross-algorithm generalization.