ACE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments
AI 摘要
ACE-Bench通过轻量级环境和可控难度评估AI Agent的推理能力。
主要贡献
- 提出 ACE-Bench 基准测试,解决了现有基准测试的局限性
- 设计了统一的基于网格的规划任务,具有可扩展的范围和可控的难度
- 提供轻量级的评估环境,提高了效率和可重复性
方法论
构建网格规划任务,通过控制隐藏槽的数量和诱饵预算来调节任务的范围和难度,并使用静态 JSON 文件模拟环境互动。
原文摘要
Existing Agent benchmarks suffer from two critical limitations: high environment interaction overhead (up to 41\% of total evaluation time) and imbalanced task horizon and difficulty distributions that make aggregate scores unreliable. To address these issues, we propose ACE-Bench built around a unified grid-based planning task, where agents must fill hidden slots in a partially completed schedule subject to both local slot constraints and global constraints. Our benchmark offers fine-grained control through two orthogonal axes: Scalable Horizons, controlled by the number of hidden slots $H$, and Controllable Difficulty, governed by a decoy budget $B$ that determines the number of globally misleading decoy candidates. Crucially, all tool calls are resolved via static JSON files under a Lightweight Environment design, eliminating setup overhead and enabling fast, reproducible evaluation suitable for training-time validation. We first validate that H and B provide reliable control over task horizon and difficulty, and that ACE-Bench exhibits strong domain consistency and model discriminability. We then conduct comprehensive experiments across 13 models of diverse sizes and families over 6 domains, revealing significant cross-model performance variation and confirming that ACE-Bench provides interpretable and controllable evaluation of agent reasoning.