Multimodal Learning 相关度: 9/10

Multimodal Backdoor Attack on VLMs for Autonomous Driving via Graffiti and Cross-Lingual Triggers

Jiancheng Wang, Lidan Liang, Yong Wang, Zengzhen Su, Haifeng Xia, Yuanting Yan, Wei Wang
arXiv: 2604.04630v1 发布: 2026-04-06 更新: 2026-04-06

AI 摘要

提出一种针对自动驾驶VLM的隐蔽多模态后门攻击,利用涂鸦和跨语言触发器,难以检测。

主要贡献

  • 提出基于涂鸦的视觉后门触发
  • 提出基于跨语言的文本后门触发
  • 证明该攻击在DriveVLM上的有效性和隐蔽性

方法论

使用stable diffusion inpainting生成自然涂鸦,结合跨语言文本触发器,通过中毒数据训练VLM,实现后门攻击。

原文摘要

Visual language model (VLM) is rapidly being integrated into safety-critical systems such as autonomous driving, making it an important attack surface for potential backdoor attacks. Existing backdoor attacks mainly rely on unimodal, explicit, and easily detectable triggers, making it difficult to construct both covert and stable attack channels in autonomous driving scenarios. GLA introduces two naturalistic triggers: graffiti-based visual patterns generated via stable diffusion inpainting, which seamlessly blend into urban scenes, and cross-language text triggers, which introduce distributional shifts while maintaining semantic consistency to build robust language-side trigger signals. Experiments on DriveVLM show that GLA requires only a 10\% poisoning ratio to achieve a 90\% Attack Success Rate (ASR) and a 0\% False Positive Rate (FPR). More insidiously, the backdoor does not weaken the model on clean tasks, but instead improves metrics such as BLEU-1, making it difficult for traditional performance-degradation-based detection methods to identify the attack. This study reveals underestimated security threats in self-driving VLMs and provides a new attack paradigm for backdoor evaluation in safety-critical multimodal systems.

标签

后门攻击 多模态学习 自动驾驶 视觉语言模型 对抗性攻击

arXiv 分类

cs.CV