Multimodal Learning 相关度: 9/10

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation

Wenli Zhang, Xianglong Shi, Sirui Zhao, Xinqi Chen, Guo Cheng, Yifan Xu, Tong Xu, Yong Liao
arXiv: 2604.08405v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

SyncBreaker提出了一种多模态对抗攻击框架,旨在破坏语音驱动的头部生成模型的同步性。

主要贡献

  • 提出多间隔采样的空化监督,引导生成静态人像
  • 提出跨注意力欺骗,抑制音频条件下的跨注意力响应
  • 提出一种用于语音驱动头部生成模型的白盒对抗攻击框架

方法论

通过联合扰动人像和音频输入,并采用多间隔采样和跨注意力欺骗来优化图像和音频流,从而破坏模型的同步性。

原文摘要

Diffusion-based audio-driven talking-head generation enables realistic portrait animation, but also introduces risks of misuse, such as fraud and misinformation. Existing protection methods are largely limited to a single modality, and neither image-only nor audio-only attacks can effectively suppress speech-driven facial dynamics. To address this gap, we propose SyncBreaker, a stage-aware multimodal protection framework that jointly perturbs portrait and audio inputs under modality-specific perceptual constraints. Our key contributions are twofold. First, for the image stream, we introduce nullifying supervision with Multi-Interval Sampling (MIS) across diffusion stages to steer the generation toward the static reference portrait by aggregating guidance from multiple denoising intervals. Second, for the audio stream, we propose Cross-Attention Fooling (CAF), which suppresses interval-specific audio-conditioned cross-attention responses. Both streams are optimized independently and combined at inference time to enable flexible deployment. We evaluate SyncBreaker in a white-box proactive protection setting. Extensive experiments demonstrate that SyncBreaker more effectively degrades lip synchronization and facial dynamics than strong single-modality baselines, while preserving input perceptual quality and remaining robust under purification. Code: https://github.com/kitty384/SyncBreaker.

标签

对抗攻击 多模态学习 语音驱动 头部生成

arXiv 分类

cs.CV