Multimodal Learning 相关度: 9/10

SnapFlow: One-Step Action Generation for Flow-Matching VLAs via Progressive Self-Distillation

Wuyang Luan, Junhui Li, Weiguang Zhao, Wenjian Zhang, Tieru Wu, Rui Ma
arXiv: 2604.05656v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

SnapFlow通过自蒸馏将多步去噪压缩为单步,加速Flow-Matching VLA模型推理。

主要贡献

  • 提出了SnapFlow,一种加速Flow-Matching VLA模型推理的自蒸馏方法
  • 通过混合标准flow-matching样本和一致性样本,避免了轨迹漂移
  • 在pi0.5和SmolVLA上验证了SnapFlow的有效性和加速效果

方法论

SnapFlow使用自蒸馏,混合flow-matching样本和两步Euler shortcut velocities计算的一致性样本,训练单步模型,避免轨迹漂移。

原文摘要

Vision-Language-Action (VLA) models based on flow matching -- such as pi0, pi0.5, and SmolVLA -- achieve state-of-the-art generalist robotic manipulation, yet their iterative denoising, typically 10 ODE steps, introduces substantial latency: on a modern GPU, denoising alone accounts for 80% of end-to-end inference time. Naively reducing the step count is unreliable, degrading success on most tasks due to the velocity field being uncalibrated for single-step jumps. We present SnapFlow, a plug-and-play self-distillation method that compresses multi-step denoising into a single forward pass (1-NFE) for flow-matching VLAs. SnapFlow mixes standard flow-matching samples with consistency samples whose targets are two-step Euler shortcut velocities computed from the model's own marginal velocity predictions, avoiding the trajectory drift caused by conditional velocities, as we analyze theoretically. A zero-initialized target-time embedding lets the network switch between local velocity estimation and global one-step generation within a single architecture. SnapFlow requires no external teacher, no architecture changes, and trains in ~12h on a single GPU. We validate on two VLA architectures spanning a 6x parameter range, with identical hyperparameters: on pi0.5 (3B) across four LIBERO suites (40 tasks, 400 episodes), SnapFlow achieves 98.75% average success -- matching the 10-step teacher at 97.75% and slightly exceeding it -- with 9.6x denoising speedup and end-to-end latency reduced from 274ms to 83ms; on SmolVLA (500M), it reduces MSE by 8.3% with 3.56x end-to-end acceleration. An action-step sweep on long-horizon tasks reveals that SnapFlow maintains its advantage across execution horizons, achieving 93% at n_act=5 where the baseline reaches only 90%. SnapFlow is orthogonal to layer-distillation and token-pruning approaches, enabling compositional speedups.

标签

Flow Matching Vision-Language-Action Self-Distillation Robotics Model Compression

arXiv 分类

cs.CV cs.AI