Multimodal Learning 相关度: 9/10

Uncovering Linguistic Fragility in Vision-Language-Action Models via Diversity-Aware Red Teaming

Baoshun Tong, Haoran He, Ling Pan, Yang Liu, Liang Lin
arXiv: 2604.05595v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

提出DAERT框架,通过多样性红队测试发现VLA模型在语言理解上的脆弱性,提高其安全性。

主要贡献

  • 提出 Diversity-Aware Embodied Red Teaming (DAERT) 框架
  • 利用均匀策略生成多样且有效的对抗性指令
  • 实验证明DAERT能有效降低VLA模型的任务成功率

方法论

设计一种基于评估均匀策略的红队测试框架,生成多样性的对抗性指令,并通过物理模拟评估VLA模型的脆弱性。

原文摘要

Vision-Language-Action (VLA) models have achieved remarkable success in robotic manipulation. However, their robustness to linguistic nuances remains a critical, under-explored safety concern, posing a significant safety risk to real-world deployment. Red teaming, or identifying environmental scenarios that elicit catastrophic behaviors, is an important step in ensuring the safe deployment of embodied AI agents. Reinforcement learning (RL) has emerged as a promising approach in automated red teaming that aims to uncover these vulnerabilities. However, standard RL-based adversaries often suffer from severe mode collapse due to their reward-maximizing nature, which tends to converge to a narrow set of trivial or repetitive failure patterns, failing to reveal the comprehensive landscape of meaningful risks. To bridge this gap, we propose a novel \textbf{D}iversity-\textbf{A}ware \textbf{E}mbodied \textbf{R}ed \textbf{T}eaming (\textbf{DAERT}) framework, to expose the vulnerabilities of VLAs against linguistic variations. Our design is based on evaluating a uniform policy, which is able to generate a diverse set of challenging instructions while ensuring its attack effectiveness, measured by execution failures in a physical simulator. We conduct extensive experiments across different robotic benchmarks against two state-of-the-art VLAs, including $π_0$ and OpenVLA. Our method consistently discovers a wider range of more effective adversarial instructions that reduce the average task success rate from 93.33\% to 5.85\%, demonstrating a scalable approach to stress-testing VLA agents and exposing critical safety blind spots before real-world deployment.

标签

Vision-Language-Action Models Red Teaming Diversity Robotics

arXiv 分类

cs.RO cs.CV