CodeTracer: Towards Traceable Agent States
AI 摘要
CodeTracer提出了一种可追踪Agent状态的架构,用于调试复杂代码Agent,并实现了错误定位。
主要贡献
- 提出了CodeTracer架构,用于解析Agent运行状态并构建追踪树。
- 构建了CodeTraceBench,一个包含大量Agent执行轨迹的基准数据集。
- 实验证明CodeTracer优于现有方法,并能通过诊断信号恢复失败的Agent运行。
方法论
通过演进的提取器解析运行工件,重构状态转换历史为分层追踪树,并定位错误源及其下游链。
原文摘要
Code agents are advancing rapidly, but debugging them is becoming increasingly difficult. As frameworks orchestrate parallel tool calls and multi-stage workflows over complex tasks, making the agent's state transitions and error propagation hard to observe. In these runs, an early misstep can trap the agent in unproductive loops or even cascade into fundamental errors, forming hidden error chains that make it hard to tell when the agent goes off track and why. Existing agent tracing analyses either focus on simple interaction or rely on small-scale manual inspection, which limits their scalability and usefulness for real coding workflows. We present CodeTracer, a tracing architecture that parses heterogeneous run artifacts through evolving extractors, reconstructs the full state transition history as a hierarchical trace tree with persistent memory, and performs failure onset localization to pinpoint the failure origin and its downstream chain. To enable systematic evaluation, we construct CodeTraceBench from a large collection of executed trajectories generated by four widely used code agent frameworks on diverse code tasks (e.g., bug fixing, refactoring, and terminal interaction), with supervision at both the stage and step levels for failure localization. Experiments show that CodeTracer substantially outperforms direct prompting and lightweight baselines, and that replaying its diagnostic signals consistently recovers originally failed runs under matched budgets. Our code and data are publicly available.