Multimodal Learning 相关度: 9/10

Challenging Vision-Language Models with Physically Deployable Multimodal Semantic Lighting Attacks

Yingying Zhao, Chengyin Hu, Qike Zhang, Xin Li, Xin Wang, Yiwei Wei, Jiujiang Guo, Jiahuan Long, Tingsong Jiang, Wen Yao
arXiv: 2604.12833v1 发布: 2026-04-14 更新: 2026-04-14

AI 摘要

提出了针对视觉语言模型的物理可部署多模态语义照明攻击,揭示了其在物理世界中的脆弱性。

主要贡献

  • 提出了一种新的物理可部署攻击框架MSLA
  • 揭示了VLMs对物理世界语义攻击的脆弱性
  • 在数字和物理环境中验证了MSLA的有效性和可迁移性

方法论

通过控制对抗性照明扰乱场景的语义理解,攻击语义对齐而非特定任务输出,从而误导VLMs。

原文摘要

Vision-Language Models (VLMs) have shown remarkable performance, yet their security remains insufficiently understood. Existing adversarial studies focus almost exclusively on the digital setting, leaving physical-world threats largely unexplored. As VLMs are increasingly deployed in real environments, this gap becomes critical, since adversarial perturbations must be physically realizable. Despite this practical relevance, physical attacks against VLMs have not been systematically studied. Such attacks may induce recognition failures and further disrupt multimodal reasoning, leading to severe semantic misinterpretation in downstream tasks. Therefore, investigating physical attacks on VLMs is essential for assessing their real-world security risks. To address this gap, we propose Multimodal Semantic Lighting Attacks (MSLA), the first physically deployable adversarial attack framework against VLMs. MSLA uses controllable adversarial lighting to disrupt multimodal semantic understanding in real scenes, attacking semantic alignment rather than only task-specific outputs. Consequently, it degrades zero-shot classification performance of mainstream CLIP variants while inducing severe semantic hallucinations in advanced VLMs such as LLaVA and BLIP across image captioning and visual question answering (VQA). Extensive experiments in both digital and physical domains demonstrate that MSLA is effective, transferable, and practically realizable. Our findings provide the first evidence that VLMs are highly vulnerable to physically deployable semantic attacks, exposing a previously overlooked robustness gap and underscoring the urgent need for physical-world robustness evaluation of VLMs.

标签

Vision-Language Models Adversarial Attack Physical Attack Multimodal Semantic Attack

arXiv 分类

cs.CV