AI Agents 相关度: 8/10

Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution

Kai Yu, Zhenhao Zhou, Junhao Zeng, Ying Wang, Xueying Du, Zhiqiang Yuan, Junwei Liu, Ziyu Zhou, Yujia Wang, Chong Wang, Xin Peng
arXiv: 2604.05955v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

该论文提出设计感知的代码修复评估基准,发现现有LLM在修复问题时设计约束不达标。

主要贡献

  • 提出了设计感知的代码修复问题,关注隐式设计约束
  • 构建了包含495个问题和1787个验证设计约束的基准ench{}
  • 实验表明基于测试的正确性高估了补丁质量,且功能正确性与设计满足感关联性低

方法论

通过挖掘真实Pull Request中的设计约束,构建基准数据集,并使用LLM验证补丁是否符合设计约束。

原文摘要

Repository-level issue resolution benchmarks have become a standard testbed for evaluating LLM-based agents, yet success is still predominantly measured by test pass rates. In practice, however, acceptable patches must also comply with project-specific design constraints, such as architectural conventions, error-handling policies, and maintainability requirements, which are rarely encoded in tests and are often documented only implicitly in code review discussions. This paper introduces \textit{design-aware issue resolution} and presents \bench{}, a benchmark that makes such implicit design constraints explicit and measurable. \bench{} is constructed by mining and validating design constraints from real-world pull requests, linking them to issue instances, and automatically checking patch compliance using an LLM-based verifier, yielding 495 issues and 1,787 validated constraints across six repositories, aligned with SWE-bench-Verified and SWE-bench-Pro. Experiments with state-of-the-art agents show that test-based correctness substantially overestimates patch quality: fewer than half of resolved issues are fully design-satisfying, design violations are widespread, and functional correctness exhibits negligible statistical association with design satisfaction. While providing issue-specific design guidance reduces violations, substantial non-compliance remains, highlighting a fundamental gap in current agent capabilities and motivating design-aware evaluation beyond functional correctness.

标签

LLM Code Repair Design Constraints Benchmark

arXiv 分类

cs.SE cs.AI