LLM Reasoning 相关度: 9/10

OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning

Xinyu Ma, Mingzhou Xu, Xuebo Liu, Chang Jin, Qiang Wang, Derek F. Wong, Min Zhang
arXiv: 2604.18530v1 发布: 2026-04-20 更新: 2026-04-20

AI 摘要

OGER通过离线指导和在线探索相结合,提升LLM在推理任务上的性能和泛化能力。

主要贡献

  • 提出OGER框架,结合离线指导和在线强化学习
  • 设计了基于熵的辅助探索奖励,鼓励模型自主探索
  • 实验证明OGER在数学和通用推理任务上优于现有方法

方法论

OGER利用多教师协同训练,构建辅助探索奖励,结合离线轨迹和模型自身熵来激励自主探索。

原文摘要

Recent advancements in Reinforcement Learning with Verifiable Rewards (RLVR) have significantly improved Large Language Model (LLM) reasoning, yet models often struggle to explore novel trajectories beyond their initial latent space. While offline teacher guidance and entropy-driven strategies have been proposed to address this, they often lack deep integration or are constrained by the model's inherent capacity. In this paper, we propose OGER, a novel framework that unifies offline teacher guidance and online reinforcement learning through a specialized reward modeling lens. OGER employs multi-teacher collaborative training and constructs an auxiliary exploration reward that leverages both offline trajectories and the model's own entropy to incentivize autonomous exploration. Extensive experiments across mathematical and general reasoning benchmarks demonstrate that OGER significantly outperforms competitive baselines, achieving substantial gains in mathematical reasoning while maintaining robust generalization to out-of-domain tasks. We provide a comprehensive analysis of training dynamics and conduct detailed ablation studies to validate the effectiveness of our entropy-aware reward modulation. Our code is available at https://github.com/ecoli-hit/OGER.git.

标签

强化学习 LLM推理 离线指导 探索奖励 熵

arXiv 分类

cs.AI