Agent Tuning & Optimization 相关度: 9/10

Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling

Jiaxuan Wang, Yulan Hu, Wenjin Yang, Zheng Pan, Xin Li, Lan-Zhe Guo
arXiv: 2604.08178v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

提出了Plan-RewardBench,用于评估智能体环境中轨迹级奖励模型的性能,揭示现有模型在长序列规划中的不足。

主要贡献

  • 提出了Plan-RewardBench基准测试,用于评估智能体轨迹级奖励模型。
  • 涵盖安全拒绝、工具无关性、复杂规划和错误恢复等四个任务。
  • 分析了现有奖励模型在长序列规划中的失败模式。

方法论

构建正负轨迹数据集,通过多模型生成、规则扰动和LLM编辑,使用pairwise协议评估不同类型的奖励模型,并进行错误诊断。

原文摘要

In classical Reinforcement Learning from Human Feedback (RLHF), Reward Models (RMs) serve as the fundamental signal provider for model alignment. As Large Language Models evolve into agentic systems capable of autonomous tool invocation and complex reasoning, the paradigm of reward modeling faces unprecedented challenges--most notably, the lack of benchmarks specifically designed to assess RM capabilities within tool-integrated environments. To address this gap, we present Plan-RewardBench, a trajectory-level preference benchmark designed to evaluate how well judges distinguish preferred versus distractor agent trajectories in complex tool-using scenarios. Plan-RewardBench covers four representative task families -- (i) Safety Refusal, (ii) Tool-Irrelevance / Unavailability, (iii) Complex Planning, and (iv) Robust Error Recovery -- comprising validated positive trajectories and confusable hard negatives constructed via multi-model natural rollouts, rule-based perturbations, and minimal-edit LLM perturbations. We benchmark representative RMs (generative, discriminative, and LLM-as-Judge) under a unified pairwise protocol, reporting accuracy trends across varying trajectory lengths and task categories. Furthermore, we provide diagnostic analyses of prevalent failure modes. Our results reveal that all three evaluator families face substantial challenges, with performance degrading sharply on long-horizon trajectories, underscoring the necessity for specialized training in agentic, trajectory-level reward modeling. Ultimately, Plan-RewardBench aims to serve as both a practical evaluation suite and a reusable blueprint for constructing agentic planning preference data.

标签

RLHF Reward Modeling AI Agents Benchmark Planning

arXiv 分类

cs.AI