Multimodal Learning 相关度: 8/10

Listening Deepfake Detection: A New Perspective Beyond Speaking-Centric Forgery Analysis

Miao Liu, Fangda Wei, Jing Wang, Xinyuan Qian
arXiv: 2604.12650v1 发布: 2026-04-14 更新: 2026-04-14

AI 摘要

论文提出听觉深度伪造检测任务,构建数据集ListenForge,并提出MANet模型,用于检测听觉伪造视频。

主要贡献

  • 提出了听觉深度伪造检测(LDD)任务,拓展了深度伪造检测的研究方向。
  • 构建了首个专门用于听觉深度伪造检测的数据集ListenForge。
  • 提出了MANet模型,利用运动信息和音频语义进行听觉伪造检测,性能优于现有模型。

方法论

提出了一个Motion-aware和Audio-guided网络(MANet),该网络捕获听者视频中的微妙运动不一致性,并利用说话者的音频语义来指导跨模态融合。

原文摘要

Existing deepfake detection research has primarily focused on scenarios where the manipulated subject is actively speaking, i.e., generating fabricated content by altering the speaker's appearance or voice. However, in realistic interaction settings, attackers often alternate between falsifying speaking and listening states to mislead their targets, thereby enhancing the realism and persuasiveness of the scenario. Although the detection of 'listening deepfakes' remains largely unexplored and is hindered by a scarcity of both datasets and methodologies, the relatively limited quality of synthesized listening reactions presents an excellent breakthrough opportunity for current deepfake detection efforts. In this paper, we present the task of Listening Deepfake Detection (LDD). We introduce ListenForge, the first dataset specifically designed for this task, constructed using five Listening Head Generation (LHG) methods. To address the distinctive characteristics of listening forgeries, we propose MANet, a Motion-aware and Audio-guided Network that captures subtle motion inconsistencies in listener videos while leveraging speaker's audio semantics to guide cross-modal fusion. Extensive experiments demonstrate that existing Speaking Deepfake Detection (SDD) models perform poorly in listening scenarios. In contrast, MANet achieves significantly superior performance on ListenForge. Our work highlights the necessity of rethinking deepfake detection beyond the traditional speaking-centric paradigm and opens new directions for multimodal forgery analysis in interactive communication settings. The dataset and code are available at https://anonymous.4open.science/r/LDD-B4CB.

标签

Deepfake Detection Multimodal Learning Audio-Visual Analysis

arXiv 分类

cs.CV cs.MM