Agent Tuning & Optimization 相关度: 9/10

Escaping the Context Bottleneck: Active Context Curation for LLM Agents via Reinforcement Learning

Xiaozhe Li, Tianyi Lyu, Yizhao Yang, Liang Shan, Siyi Yang, Ligao Zhang, Zhuoyi Huang, Qingwen Liu, Yang Li
arXiv: 2604.11462v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

提出ContextCurator框架,通过强化学习主动管理LLM上下文,提升长程任务表现。

主要贡献

  • 提出ContextCurator框架,解耦上下文管理与任务执行
  • 使用强化学习训练ContextCurator,主动过滤噪声信息
  • 实验证明ContextCurator在长程任务中能提升性能并减少token消耗

方法论

使用轻量级策略模型ContextCurator,通过强化学习训练,动态选择和保留上下文信息,减少token消耗。

原文摘要

Large Language Models (LLMs) struggle with long-horizon tasks due to the "context bottleneck" and the "lost-in-the-middle" phenomenon, where accumulated noise from verbose environments degrades reasoning over multi-turn interactions. To address this issue, we introduce a symbiotic framework that decouples context management from task execution. Our architecture pairs a lightweight, specialized policy model, ContextCurator, with a powerful frozen foundation model, TaskExecutor. Trained via reinforcement learning, ContextCurator actively reduces information entropy in the working memory. It aggressively prunes environmental noise while preserving reasoning anchors, that is, sparse data points that are critical for future deductions. On WebArena, our framework improves the success rate of Gemini-3.0-flash from 36.4% to 41.2% while reducing token consumption by 8.8% (from 47.4K to 43.3K). On DeepSearch, it achieves a 57.1% success rate, compared with 53.9%, while reducing token consumption by a factor of 8. Remarkably, a 7B ContextCurator matches the context management performance of GPT-4o, providing a scalable and computationally efficient paradigm for autonomous long-horizon agents.

标签

AI Agents Reinforcement Learning Context Management LLM Optimization

arXiv 分类

cs.AI