AI Agents 相关度: 9/10

Value-Guidance MeanFlow for Offline Multi-Agent Reinforcement Learning

Teng Pang, Zhiqiang Dong, Yan Zhang, Rongjian Xu, Guoqiang Wu, Yilong Yin
arXiv: 2604.08174v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

提出Value-Guidance MeanFlow策略(VGM$^2$P),高效解决离线多智能体强化学习的分布偏移和效率问题。

主要贡献

  • 提出Value Guidance MeanFlow Policy (VGM$^2$P)
  • 利用全局优势值指导智能体协作,进行条件行为克隆
  • 使用classifier-free guidance MeanFlow提高策略表达性和推理效率

方法论

通过全局优势值引导,将最优策略学习视为条件行为克隆,并利用MeanFlow提升效率。

原文摘要

Offline multi-agent reinforcement learning (MARL) aims to learn the optimal joint policy from pre-collected datasets, requiring a trade-off between maximizing global returns and mitigating distribution shift from offline data. Recent studies use diffusion or flow generative models to capture complex joint policy behaviors among agents; however, they typically rely on multi-step iterative sampling, thereby reducing training and inference efficiency. Although further research improves sampling efficiency through methods like distillation, it remains sensitive to the behavior regularization coefficient. To address the above-mentioned issues, we propose Value Guidance Multi-agent MeanFlow Policy (VGM$^2$P), a simple yet effective flow-based policy learning framework that enables efficient action generation with coefficient-insensitive conditional behavior cloning. Specifically, VGM$^2$P uses global advantage values to guide agent collaboration, treating optimal policy learning as conditional behavior cloning. Additionally, to improve policy expressiveness and inference efficiency in multi-agent scenarios, it leverages classifier-free guidance MeanFlow for both policy training and execution. Experiments on tasks with both discrete and continuous action spaces demonstrate that, even when trained solely via conditional behavior cloning, VGM$^2$P efficiently achieves performance comparable to state-of-the-art methods.

标签

离线多智能体强化学习 Flow模型 条件行为克隆 MeanFlow

arXiv 分类

cs.LG