Multimodal Learning 相关度: 9/10

Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs

Jianhao Chen, Haoyang Chen, Hanjie Zhao, Haozhe Liang, Tieyun Qian
arXiv: 2604.12616v1 发布: 2026-04-14 更新: 2026-04-14

AI 摘要

提出MemJack框架,利用视觉语义进行多模态大模型的越狱攻击,并发布大规模攻击数据集MemJack-Bench。

主要贡献

  • 提出了MemJack框架,用于对VLMs进行语义层面的越狱攻击。
  • 利用多智能体协作,动态映射视觉实体到恶意意图,生成对抗性提示。
  • 发布了包含超过113,000条攻击轨迹的MemJack-Bench数据集。

方法论

MemJack利用记忆增强的多智能体协作,通过视觉语义伪装生成对抗性提示,并通过几何滤波器绕过潜在空间拒绝。

原文摘要

The rapid evolution of Vision-Language Models (VLMs) has catalyzed unprecedented capabilities in artificial intelligence; however, this continuous modal expansion has inadvertently exposed a vastly broadened and unconstrained adversarial attack surface. Current multimodal jailbreak strategies primarily focus on surface-level pixel perturbations and typographic attacks or harmful images; however, they fail to engage with the complex semantic structures intrinsic to visual data. This leaves the vast semantic attack surface of original, natural images largely unscrutinized. Driven by the need to expose these deep-seated semantic vulnerabilities, we introduce \textbf{MemJack}, a \textbf{MEM}ory-augmented multi-agent \textbf{JA}ilbreak atta\textbf{CK} framework that explicitly leverages visual semantics to orchestrate automated jailbreak attacks. MemJack employs coordinated multi-agent cooperation to dynamically map visual entities to malicious intents, generate adversarial prompts via multi-angle visual-semantic camouflage, and utilize an Iterative Nullspace Projection (INLP) geometric filter to bypass premature latent space refusals. By accumulating and transferring successful strategies through a persistent Multimodal Experience Memory, MemJack maintains highly coherent extended multi-turn jailbreak attack interactions across different images, thereby improving the attack success rate (ASR) on new images. Extensive empirical evaluations across full, unmodified COCO val2017 images demonstrate that MemJack achieves a 71.48\% ASR against Qwen3-VL-Plus, scaling to 90\% under extended budgets. Furthermore, to catalyze future defensive alignment research, we will release \textbf{MemJack-Bench}, a comprehensive dataset comprising over 113,000 interactive multimodal jailbreak attack trajectories, establishing a vital foundation for developing inherently robust VLMs.

标签

multimodal jailbreak adversarial attack VLM multi-agent

arXiv 分类

cs.AI cs.MM