Agent Tuning & Optimization 相关度: 9/10

Select Smarter, Not More: Prompt-Aware Evaluation Scheduling with Submodular Guarantees

Xiaoyu Ma, Yiwen Li, Haoyue Liu, Zhichao Wang, Ye Chen, Yongxin Guo, Xiaoying Tang
arXiv: 2604.11328v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

提出Prompt-Aware Online Evaluation Scheduling (POES),提升自动prompt优化效率和效果。

主要贡献

  • 提出了POES框架,通过在线自适应评估调度提升prompt优化效果
  • 结合IRT、覆盖率和切换成本,构建了单调次模目标函数
  • 实验证明POES在多个任务上显著提升了prompt优化精度和效率

方法论

POES利用在线自适应测试思路,结合IRT等理论构建优化目标,贪婪选择评估样本,平衡探索与利用。

原文摘要

Automatic prompt optimization (APO) hinges on the quality of its evaluation signal, yet scoring every prompt candidate on the full training set is prohibitively expensive. Existing methods either fix a single evaluation subset before optimization begins (principled but prompt-agnostic) or adapt it heuristically during optimization (flexible but unstable and lacking formal guarantees). We observe that APO naturally maps to an online adaptive testing problem: prompts are examinees, training examples are test items, and the scheduler should select items that best discriminate among the strongest candidates. This insight motivates Prompt-Aware Online Evaluation Scheduling (POES), which integrates an IRT-based discrimination utility, a facility-location coverage term, and switching-cost-aware warm-start swaps into a unified objective that is provably monotone submodular, yielding a (1-1/e) greedy guarantee for cold starts and bounded drift for warm-start updates. An adaptive controller modulates the exploration-exploitation balance based on optimization progress. Across 36 tasks spanning three benchmark families, POES achieves the highest overall average accuracy (6.2 percent improvement over the best baseline) with negligible token overhead (approximately 4 percent) at the same evaluation budget. Moreover, principled selection at k = 20 examples matches or exceeds the performance of naive evaluation at k = 30-50, reducing token consumption by 35-60 percent, showing that selecting smarter is more effective than selecting more. Our results demonstrate that evaluation scheduling is a first-class component of APO, not an implementation detail.

标签

prompt optimization evaluation scheduling submodular optimization adaptive testing

arXiv 分类

cs.AI cs.LG