LLM Reasoning 相关度: 9/10

DiPO: Disentangled Perplexity Policy Optimization for Fine-grained Exploration-Exploitation Trade-Off

Xiaofan Li, Ming Yang, Zhiyuan Ma, Shichao Ma, Jintao Du, Yu Cheng, Weiqiang Wang, Zhizhong Zhang, Xin Tan, Yanyun Qu, Lizhuang Ma, Yuan Xie
arXiv: 2604.13902v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

DiPO通过解耦困惑度空间实现细粒度的探索-利用权衡,提升LLM推理能力。

主要贡献

  • 提出基于困惑度空间解耦的探索-利用策略
  • 设计双向奖励分配机制,稳定策略优化
  • 在数学推理和函数调用任务上验证了方法有效性

方法论

通过困惑度将样本空间划分为探索和利用区域,并利用双向奖励机制引导LLM进行更有效的策略优化。

原文摘要

Reinforcement Learning with Verifiable Rewards (RLVR) has catalyzed significant advances in the reasoning capabilities of Large Language Models (LLMs). However, effectively managing the exploration and exploitation trade-off remains a critical challenge. In this paper, we fully analyze the exploration and exploitation dilemma of extremely hard and easy samples during the training and propose a new fine-grained trade-off mechanism. Concretely, we introduce a perplexity space disentangling strategy that divides the sample space into distinct exploration (high perplexity) and exploitation (low perplexity) subspaces, thereby mining fine-grained samples requiring exploration-exploitation trade-off. Subsequently, we propose a bidirectional reward allocation mechanism with a minimum impact on verification rewards to implement perplexity-guided exploration and exploitation, enabling more stable policy optimization. Finally, we have evaluated our method on two mainstream tasks: mathematical reasoning and function calling, and experimental results demonstrate the superiority of the proposed method, confirming its effectiveness in enhancing LLM performance by fine-grained exploration-exploitation trade-off.

标签

强化学习 LLM 探索-利用 困惑度 奖励机制

arXiv 分类

cs.LG