AI Agents 相关度: 9/10

$π$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data

Yaocheng Zhang, Yuanheng Zhu, Wenyue Chong, Songjun Tu, Qichao Zhang, Jiajun Chai, Xiaohan Wang, Wei Lin, Guojun Yin, Dongbin Zhao
arXiv: 2604.14054v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

提出了一种基于特权信息自蒸馏的多智能体自博弈框架,显著提升了搜索agent的训练效率。

主要贡献

  • 提出 Privileged Information Self-Play ($π$-Play) 框架
  • 利用任务生成过程中的问题构造路径 (QCP) 作为特权信息
  • 将稀疏奖励自博弈转化为密集反馈自进化循环

方法论

设计一个examiner生成任务和QCP,teacher模型利用QCP指导student模型,进行自蒸馏。

原文摘要

Deep search agents have emerged as a promising paradigm for addressing complex information-seeking tasks, but their training remains challenging due to sparse rewards, weak credit assignment, and limited labeled data. Self-play offers a scalable route to reduce data dependence, but conventional self-play optimizes students only through sparse outcome rewards, leading to low learning efficiency. In this work, we observe that self-play naturally produces a question construction path (QCP) during task generation, an intermediate artifact that captures the reverse solution process. This reveals a new source of privileged information for self-distillation: self-play can itself provide high-quality privileged context for the teacher model in a low-cost and scalable manner, without relying on human feedback or curated privileged information. Leveraging this insight, we propose Privileged Information Self-Play ($π$-Play), a multi-agent self-evolution framework. In $π$-Play, an examiner generates tasks together with their QCPs, and a teacher model leverages QCP as privileged context to densely supervise a student via self-distillation. This design transforms conventional sparse-reward self-play into a dense-feedback self-evolution loop. Extensive experiments show that data-free $π$-Play surpasses fully supervised search agents and improves evolutionary efficiency by 2-3$\times$ over conventional self-play.

标签

自博弈 自蒸馏 多智能体 搜索Agent

arXiv 分类

cs.LG cs.CL