Agent Tuning & Optimization 相关度: 7/10

Triviality Corrected Endogenous Reward

Xinda Wang, Zhengxu Hou, Yangshijie Zhang, Bingren Yan, Jialin Liu, Chenzhuo Zhao, Zhibo Yang, Bin-Bin Yang, Feng Xiao
arXiv: 2604.11522v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

提出TCER方法,通过纠正平凡偏差,提升开放式文本生成的质量和多样性,无需外部监督。

主要贡献

  • 提出了TCER方法,解决了开放式文本生成中的平凡偏差问题
  • 使用相对信息增益作为奖励信号,鼓励生成多样化内容
  • 无需外部监督,在多个文本生成任务上取得了显著提升

方法论

TCER通过奖励 specialist policy 相对于 generalist policy 的信息增益,并使用概率依赖的纠正机制来解决平凡偏差。

原文摘要

Reinforcement learning for open-ended text generation is constrained by the lack of verifiable rewards, necessitating reliance on judge models that require either annotated data or powerful closed-source models. Inspired by recent work on unsupervised reinforcement learning for mathematical reasoning using confidence-based endogenous rewards, we investigate whether this principle can be adapted to open-ended writing tasks. We find that directly applying confidence rewards leads to Triviality Bias: the policy collapses toward high-probability outputs, reducing diversity and meaningful content. We propose TCER (Triviality Corrected Endogenous Reward), which addresses this bias by rewarding the relative information gain between a specialist policy and a generalist reference policy, modulated by a probability-dependent correction mechanism. Across multiple writing benchmarks and model architectures, TCER achieves consistent improvements without external supervision. Furthermore, TCER also transfers effectively to mathematical reasoning, validating the generality of our approach across different generation tasks.

标签

强化学习 文本生成 无监督学习 奖励设计

arXiv 分类

cs.CL