Multimodal Learning 相关度: 8/10

Jump-Start Reinforcement Learning with Vision-Language-Action Regularization

Angelo Moroncelli, Roberto Zanetti, Marco Maccarini, Loris Roveda
arXiv: 2604.13733v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

VLAJS方法结合视觉-语言-动作模型和强化学习,提升机器人操作任务的探索效率和学习效果。

主要贡献

  • 提出 VLAJS 方法,利用 VLA 模型指导强化学习探索
  • 引入动作一致性正则化,软对齐 VLA 指导与 RL 策略
  • 实验验证 VLAJS 在多个机器人操作任务上的优越性

方法论

VLAJS 使用稀疏的 VLA 指导增强 PPO 算法,通过动作一致性正则化,在早期训练中引导 RL 策略,并逐渐退火。

原文摘要

Reinforcement learning (RL) enables high-frequency, closed-loop control for robotic manipulation, but scaling to long-horizon tasks with sparse or imperfect rewards remains difficult due to inefficient exploration and poor credit assignment. Vision-Language-Action (VLA) models leverage large-scale multimodal pretraining to provide generalist, task-level reasoning, but current limitations hinder their direct use in fast and precise manipulation. In this paper, we propose Vision-Language-Action Jump-Starting (VLAJS), a method that bridges sparse VLA guidance with on-policy RL to improve exploration and learning efficiency. VLAJS treats VLAs as transient sources of high-level action suggestions that bias early exploration and improve credit assignment, while preserving the high-frequency, state-based control of RL. Our approach augments Proximal Policy Optimization (PPO) with a directional action-consistency regularization that softly aligns the RL agent's actions with VLA guidance during early training, without enforcing strict imitation, requiring demonstrations, or relying on continuous teacher queries. VLA guidance is applied sparsely and annealed over time, allowing the agent to adapt online and ultimately surpass the guiding policy. We evaluate VLAJS on six challenging manipulation tasks: lifting, pick-and-place, peg reorientation, peg insertion, poking, and pushing in simulation, and validate a subset on a real Franka Panda robot. VLAJS consistently outperforms PPO and distillation-style baselines in sample efficiency, reducing required environment interactions by over 50% in several tasks. Real-world experiments demonstrate zero-shot sim-to-real transfer and robust execution under clutter, object variation, and external perturbations.

标签

强化学习 视觉语言 机器人操作

arXiv 分类

cs.LG cs.AI cs.RO