Agent Tuning & Optimization 相关度: 9/10

Utilizing and Calibrating Hindsight Process Rewards via Reinforcement with Mutual Information Self-Evaluation

Jiashu Yao, Heyan Huang, Zeming Liu, Yuhang Guo
arXiv: 2604.11611v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

MISE通过后见之明自评估和互信息校准,解决了LLM智能体强化学习中的稀疏奖励问题。

主要贡献

  • 提出了MISE方法,利用后见之明生成自评估作为密集奖励
  • 提供了生成自奖励范式的首个形式化基础
  • 通过实验验证MISE优于基线方法,使开源LLM达到媲美GPT-4o的性能

方法论

利用后见之明自评估生成密集奖励,并通过互信息校准,使其与环境反馈对齐,进行强化学习。

原文摘要

To overcome the sparse reward challenge in reinforcement learning (RL) for agents based on large language models (LLMs), we propose Mutual Information Self-Evaluation (MISE), an RL paradigm that utilizes hindsight generative self-evaluation as dense reward signals while simultaneously calibrating them against the environmental feedbacks. Empirically, MISE enables an agent to learn autonomously from dense internal rewards supplementing sparse extrinsic signals. Theoretically, our work provides the first formal foundation for the paradigm of generative self-rewarding. We prove that utilizing hindsight self-evaluation rewards is equivalent to minimizing an objective that combines mutual information with a KL divergence term between the policy and a proxy reward policy. This theoretical insight then informs and justifies our calibration step, which actively aligns these rewards with the optimal policy. Extensive experiments show that MISE outperforms strong baselines, enabling open-source LLMs about 7B parameters to achieve performance comparable to GPT-4o on validation without expert supervision.

标签

Reinforcement Learning Large Language Models Self-Evaluation Mutual Information

arXiv 分类

cs.CL cs.LG