Agent Tuning & Optimization 相关度: 9/10

Does RL Expand the Capability Boundary of LLM Agents? A PASS@(k,T) Analysis

Zhiyuan Zhai, Wenjing Yan, Xiaodan Shao, Xin Wang
arXiv: 2604.14877v1 发布: 2026-04-16 更新: 2026-04-16

AI 摘要

研究表明,强化学习能显著扩展LLM Agent在复杂工具使用任务上的能力边界。

主要贡献

  • 提出了PASS@(k,T)指标,用于评估RL对LLM Agent能力的扩展
  • 证明了RL能够真正扩展LLM Agent在组合工具使用任务中的能力
  • 分析了RL改善LLM Agent性能的机制,发现其通过重加权策略分布来优化信息整合

方法论

通过引入PASS@(k,T)指标,对比base模型和RL Agent在不同k和T下的表现,并进行机制分析。

原文摘要

Does reinforcement learning genuinely expand what LLM agents can do, or merely make them more reliable? For static reasoning, recent work answers the second: base and RL pass@k curves converge at large k. We ask whether this holds for agentic tool use, where T rounds of interaction enable compositional strategies that re-sampling cannot recover. We introduce PASS@(k,T), a two-dimensional metric that jointly varies sampling budget k and interaction depth T, separating capability expansion from efficiency improvement. Our main finding is that, contrary to the static-reasoning result, tool-use RL genuinely enlarges the capability boundary: the RL agent's pass-curve pulls above the base model's and the gap widens at large k rather than converging. The expansion is specific to compositional, sequential information gathering; on simpler tasks RL behaves as prior work predicts. Under matched training data, supervised fine-tuning regresses the boundary on the same compositional tasks, isolating self-directed exploration as the causal factor. Mechanism analysis shows RL reweights the base strategy distribution toward the subset whose downstream reasoning more often yields a correct answer, with the improvement concentrated on how the agent integrates retrieved information. These results reconcile optimistic and pessimistic readings of RL for LLMs: both are correct, on different task types.

标签

AI Agents Reinforcement Learning LLM Tool Use Reasoning

arXiv 分类

cs.LG