Agent Tuning & Optimization 相关度: 8/10

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning

Shiwan Zhao, Zhihu Wang, Xuyang Zhao, Jiaming Zhou, Caiyue Xu, Chenfei Liu, Liting Zhang, Yuhang Jia, Yanzhe Zhang, Hualong Yu, Zichen Xu, Qicheng Li, Yong Qin
arXiv: 2604.07941v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

论文从轨迹来源角度统一分析LLM后训练方法,提出支持扩展、策略重塑和行为巩固三个核心角色。

主要贡献

  • 统一了LLM后训练的各种方法
  • 提出了基于轨迹来源的分类框架
  • 强调了系统设计在后训练中的重要性

方法论

通过分析不同后训练方法所使用的轨迹来源(on-policy vs off-policy),并结合支持扩展、策略重塑和行为巩固三个角色来理解。

原文摘要

Post-training has become central to turning pretrained large language models (LLMs) into aligned and deployable systems. Recent progress spans supervised fine-tuning (SFT), preference optimization, reinforcement learning (RL), process supervision, verifier-guided methods, distillation, and multi-stage pipelines. Yet these methods are often discussed in fragmented ways, organized by labels or objective families rather than by the behavioral bottlenecks they address. This survey argues that LLM post-training is best understood as structured intervention on model behavior. We organize the field first by trajectory provenance, which defines two primary learning regimes: off-policy learning on externally supplied trajectories, and on-policy learning on learner-generated rollouts. We then interpret methods through two recurring roles -- effective support expansion, which makes useful behaviors more reachable, and policy reshaping, which improves behavior within already reachable regions -- together with a complementary systems-level role, behavioral consolidation, which preserves, transfers, and amortizes behavior across stages and model transitions. This perspective yields a unified reading of major paradigms. SFT may serve either support expansion or policy reshaping, whereas preference-based methods are usually off-policy reshaping. On-policy RL often improves behavior on learner-generated states, though under stronger guidance it can also make hard-to-reach reasoning paths reachable. Distillation is often best understood as consolidation rather than only compression, and hybrid pipelines emerge as coordinated multi-stage compositions. Overall, the framework helps diagnose post-training bottlenecks and reason about stage composition, suggesting that progress in LLM post-training increasingly depends on coordinated system design rather than any single dominant objective.

标签

LLM Post-training Reinforcement Learning Supervised Fine-tuning Distillation

arXiv 分类

cs.CL cs.AI cs.LG