Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models
AI 摘要
论文分析了OPD训练中长度膨胀问题,并提出StableOPD框架,有效提升了模型在数学推理任务上的性能。
主要贡献
- 发现了OPD训练中长度膨胀的现象
- 提出了StableOPD框架,包括参考分布约束和混合蒸馏
- 在数学推理数据集上验证了StableOPD的有效性
方法论
提出StableOPD,使用参考分布约束限制长度膨胀,并结合rollout mixture蒸馏稳定训练过程,实验验证在数学推理任务上的性能提升。
原文摘要
On-policy distillation (OPD) trains student models under their own induced distribution while leveraging supervision from stronger teachers. We identify a failure mode of OPD: as training progresses, on-policy rollouts can undergo abrupt length inflation, causing truncated trajectories to dominate the training data. This truncation collapse coincides with abrupt repetition saturation and induces biased gradient signals, leading to severe training instability and sharp degradation in validation performance. We attribute this problem to the interaction between student-induced data collection and the distillation objective, which implicitly favors long and repetitive rollouts. To address this issue, we propose StableOPD, a stabilized OPD framework that combines a reference-based divergence constraint with rollout mixture distillation. These together mitigate repetition-induced length inflation and further stabilize OPD training. Across multiple math reasoning datasets, our approach prevents truncation collapse, stabilizes training dynamics, and improves performance by 7.2% on average.