Agent Tuning & Optimization 相关度: 8/10

From Curiosity to Caution: Mitigating Reward Hacking for Best-of-N with Pessimism

Zhuohao Yu, Zhiwei Steven Wu, Adam Block
arXiv: 2604.04648v1 发布: 2026-04-06 更新: 2026-04-06

AI 摘要

论文提出了一种名为“caution”的悲观强化学习方法,有效缓解了Best-of-N采样中常见的奖励黑客问题。

主要贡献

  • 提出了一种名为“caution”的新方法,缓解了Best-of-N采样中的奖励黑客问题。
  • 通过实验证明caution方法在实际应用中有效。
  • 提供了在简化线性环境下的理论分析,证明caution优于标准BoN方法。

方法论

通过训练一个错误模型来预测典型响应,并利用预测误差来降低非典型响应的奖励估计,从而避免不确定性高的行为。

原文摘要

Inference-time compute scaling has emerged as a powerful paradigm for improving language model performance on a wide range of tasks, but the question of how best to use the additional compute remains open. A popular approach is BoN sampling, where N candidate responses are generated, scored according to a reward model, and the highest-scoring response is selected. While this approach can improve performance, it is vulnerable to reward hacking, where performance degrades as N increases due to the selection of responses that exploit imperfections in the reward model instead of genuinely improving generation quality. Prior attempts to mitigate reward hacking, via stronger reward models or heavy-handed distributional regularization, either fail to fully address over-optimization or are too conservative to exploit additional compute. In this work, we explore the principle of pessimism in RL, which uses lower confidence bounds on value estimates to avoid OOD actions with uncertain reward estimates. Our approach, termed as caution, can be seen as the reverse of curiosity: where curiosity rewards prediction error as a signal of novelty, caution penalizes prediction error as a signal of distributional uncertainty. Practically, caution trains an error model on typical responses and uses its prediction error to lower reward estimates for atypical ones. Our extensive empirical evaluation demonstrates that caution is a simple, computationally efficient approach that substantially mitigates reward hacking in BoN sampling. We also provide a theoretical analysis in a simplified linear setting, which shows that caution provably improves over the standard BoN approach. Together, our results not only establish caution as a practical solution to reward hacking, but also provide evidence that curiosity-based approaches can be a general OOD detection technique in LLM settings.

标签

Reward Hacking Best-of-N Sampling Pessimism Out-of-Distribution Detection

arXiv 分类

cs.LG