Multimodal Learning 相关度: 9/10

Stealthy and Adjustable Text-Guided Backdoor Attacks on Multimodal Pretrained Models

Yiyang Zhang, Chaojian Yu, Ziming Hong, Yuanjie Shao, Qinmu Peng, Tongliang Liu, Xinge You
arXiv: 2604.05809v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

提出一种隐蔽且可调的文本引导多模态预训练模型后门攻击方法,提高了攻击的实用性和隐蔽性。

主要贡献

  • 提出文本引导后门攻击(TGB),利用常见词作为触发器
  • 引入视觉对抗扰动,调节模型对文本触发器的学习
  • 实验证明TGB在CIR和VQA任务上的有效性和可调性

方法论

利用常见词作为文本触发器,并结合视觉对抗扰动来控制后门攻击的学习,实现隐蔽且可调节的攻击。

原文摘要

Multimodal pretrained models are vulnerable to backdoor attacks, yet most existing methods rely on visual or multimodal triggers, which are impractical since visually embedded triggers rarely occur in real-world data. To overcome this limitation, we propose a novel Text-Guided Backdoor (TGB) attack on multimodal pretrained models, where commonly occurring words in textual descriptions serve as backdoor triggers, significantly improving stealthiness and practicality. Furthermore, we introduce visual adversarial perturbations on poisoned samples to modulate the model's learning of textual triggers, enabling a controllable and adjustable TGB attack. Extensive experiments on downstream tasks built upon multimodal pretrained models, including Composed Image Retrieval (CIR) and Visual Question Answering (VQA), demonstrate that TGB achieves practicality and stealthiness with adjustable attack success rates across diverse realistic settings, revealing critical security vulnerabilities in multimodal pretrained models.

标签

后门攻击 多模态学习 文本引导 对抗扰动 安全性

arXiv 分类

cs.CR cs.LG