Multimodal Learning 相关度: 9/10

DA-PTQ: Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models

Siyuan Xu, Tianshi Wang, Fengling Li, Lei Zhu, Heng Tao Shen
arXiv: 2604.11572v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

DA-PTQ通过降低量化误差累积,实现了在资源受限机器人上高效部署VLA模型。

主要贡献

  • 提出Drift-Aware Post-Training Quantization (DA-PTQ)方法
  • 设计Cross-Space Representation Compensation降低跨模态失真
  • 提出Motion-Driven Mixed-Precision Allocation优化轨迹运动误差

方法论

将量化视为一个漂移感知优化问题,通过跨空间表示补偿和运动驱动混合精度分配,减少量化误差累积。

原文摘要

Vision-Language-Action models (VLAs) have demonstrated strong potential for embodied AI, yet their deployment on resource-limited robots remains challenging due to high memory and computational demands. While Post-Training Quantization (PTQ) provides an efficient solution, directly applying PTQ to VLAs often results in severe performance degradation during sequential control. We identify temporal error accumulation as a key factor, where quantization perturbations at the vision-language-to-action interface are progressively amplified, leading to kinematic drift in executed trajectories. To address this issue, we propose Drift-Aware Post-Training Quantization (DA-PTQ), which formulates quantization as a drift-aware optimization problem over sequential decision processes. DA-PTQ consists of two components: (1) Cross-Space Representation Compensation, which mitigates structured distortions between multimodal representations and action space to improve action consistency, and (2) Motion-Driven Mixed-Precision Allocation, which assigns bit-widths by minimizing trajectory-level motion errors. Extensive experiments show that DA-PTQ significantly reduces kinematic drift and achieves comparable performance to full-precision models under low-bit settings, enabling practical deployment of VLAs on resource-limited robotic platforms.

标签

Post-Training Quantization Vision-Language-Action Models Embodied AI Robotics Kinematic Drift

arXiv 分类

cs.RO cs.MM