AI Agents 相关度: 9/10

ACE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments

Wang Yang, Chaoda Song, Xinpeng Li, Debargha Ganguly, Chuang Ma, Shouren Wang, Zhihao Dou, Yuli Zhou, Vipin Chaudhary, Xiaotian Han
arXiv: 2604.06111v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

ACE-Bench通过轻量级环境和可控难度评估AI Agent的推理能力。

主要贡献

  • 提出 ACE-Bench 基准测试,解决了现有基准测试的局限性
  • 设计了统一的基于网格的规划任务,具有可扩展的范围和可控的难度
  • 提供轻量级的评估环境,提高了效率和可重复性

方法论

构建网格规划任务,通过控制隐藏槽的数量和诱饵预算来调节任务的范围和难度,并使用静态 JSON 文件模拟环境互动。

原文摘要

Existing Agent benchmarks suffer from two critical limitations: high environment interaction overhead (up to 41\% of total evaluation time) and imbalanced task horizon and difficulty distributions that make aggregate scores unreliable. To address these issues, we propose ACE-Bench built around a unified grid-based planning task, where agents must fill hidden slots in a partially completed schedule subject to both local slot constraints and global constraints. Our benchmark offers fine-grained control through two orthogonal axes: Scalable Horizons, controlled by the number of hidden slots $H$, and Controllable Difficulty, governed by a decoy budget $B$ that determines the number of globally misleading decoy candidates. Crucially, all tool calls are resolved via static JSON files under a Lightweight Environment design, eliminating setup overhead and enabling fast, reproducible evaluation suitable for training-time validation. We first validate that H and B provide reliable control over task horizon and difficulty, and that ACE-Bench exhibits strong domain consistency and model discriminability. We then conduct comprehensive experiments across 13 models of diverse sizes and families over 6 domains, revealing significant cross-model performance variation and confirming that ACE-Bench provides interpretable and controllable evaluation of agent reasoning.

标签

AI Agent 基准测试 推理 评估 轻量级环境

arXiv 分类

cs.AI cs.CL