AI Agents 相关度: 9/10

Multi-ORFT: Stable Online Reinforcement Fine-Tuning for Multi-Agent Diffusion Planning in Cooperative Driving

Haojie Bai, Aimin Li, Ruoyu Yao, Xiongwei Zhao, Tingting Zhang, Xing Zhang, Lin Gao, and Jun Ma
arXiv: 2604.11734v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

提出Multi-ORFT,通过扩散模型预训练和强化学习后训练,提升多智能体协同驾驶的安全性与效率。

主要贡献

  • 提出Multi-ORFT框架,结合扩散模型和强化学习
  • 使用自注意力、交叉注意力和AdaLN-Zero改进场景一致性
  • 结合VG-GRPO稳定在线强化学习训练

方法论

结合场景条件扩散预训练与稳定在线强化后训练,使用双层MDP和方差门控策略优化。

原文摘要

Closed-loop cooperative driving requires planners that generate realistic multimodal multi-agent trajectories while improving safety and traffic efficiency. Existing diffusion planners can model multimodal behaviors from demonstrations, but they often exhibit weak scene consistency and remain poorly aligned with closed-loop objectives; meanwhile, stable online post-training in reactive multi-agent environments remains difficult. We present Multi-ORFT, which couples scene-conditioned diffusion pre-training with stable online reinforcement post-training. In pre-training, the planner uses inter-agent self-attention, cross-attention, and AdaLN-Zero-based scene conditioning to improve scene consistency and road adherence of joint trajectories. In post-training, we formulate a two-level MDP that exposes step-wise reverse-kernel likelihoods for online optimization, and combine dense trajectory-level rewards with variance-gated group-relative policy optimization (VG-GRPO) to stabilize training. On the WOMD closed-loop benchmark, Multi-ORFT reduces collision rate from 2.04% to 1.89% and off-road rate from 1.68% to 1.36%, while increasing average speed from 8.36 to 8.61 m/s relative to the pre-trained planner, and it outperforms strong open-source baselines including SMART-large, SMART-tiny-CLSFT, and VBD on the primary safety and efficiency metrics. These results show that coupling scene-consistent denoising with stable online diffusion-policy optimization improves the reliability of closed-loop cooperative driving.

标签

多智能体 协同驾驶 扩散模型 强化学习 在线学习

arXiv 分类

cs.RO cs.AI