AI Agents 相关度: 9/10

RAD-2: Scaling Reinforcement Learning in a Generator-Discriminator Framework

Hao Gao, Shaoyu Chen, Yifan Zhu, Yuehao Song, Wenyu Liu, Qian Zhang, Xinggang Wang
arXiv: 2604.15308v1 发布: 2026-04-16 更新: 2026-04-16

AI 摘要

RAD-2提出了一种生成器-判别器框架,用于提升自动驾驶规划的稳定性和安全性。

主要贡献

  • 提出RAD-2框架,结合扩散模型和强化学习进行轨迹规划
  • 引入时间一致性相对策略优化算法,改善信用分配问题
  • 提出在线生成器优化算法,引导生成器学习高回报轨迹
  • BEV-Warp高效仿真环境加速训练

方法论

采用扩散模型生成候选轨迹,强化学习优化判别器评估轨迹质量,并利用改进的策略优化算法训练。

原文摘要

High-level autonomous driving requires motion planners capable of modeling multimodal future uncertainties while remaining robust in closed-loop interactions. Although diffusion-based planners are effective at modeling complex trajectory distributions, they often suffer from stochastic instabilities and the lack of corrective negative feedback when trained purely with imitation learning. To address these issues, we propose RAD-2, a unified generator-discriminator framework for closed-loop planning. Specifically, a diffusion-based generator is used to produce diverse trajectory candidates, while an RL-optimized discriminator reranks these candidates according to their long-term driving quality. This decoupled design avoids directly applying sparse scalar rewards to the full high-dimensional trajectory space, thereby improving optimization stability. To further enhance reinforcement learning, we introduce Temporally Consistent Group Relative Policy Optimization, which exploits temporal coherence to alleviate the credit assignment problem. In addition, we propose On-policy Generator Optimization, which converts closed-loop feedback into structured longitudinal optimization signals and progressively shifts the generator toward high-reward trajectory manifolds. To support efficient large-scale training, we introduce BEV-Warp, a high-throughput simulation environment that performs closed-loop evaluation directly in Bird's-Eye View feature space via spatial warping. RAD-2 reduces the collision rate by 56% compared with strong diffusion-based planners. Real-world deployment further demonstrates improved perceived safety and driving smoothness in complex urban traffic.

标签

自动驾驶 强化学习 扩散模型 轨迹规划

arXiv 分类

cs.CV