VRAG-DFD: Verifiable Retrieval-Augmentation for MLLM-based Deepfake Detection
AI 摘要
VRAG-DFD通过RAG和RL增强MLLM,提升Deepfake检测能力,构建了专业知识库和推理数据集。
主要贡献
- 提出了VRAG-DFD框架,结合RAG和RL
- 构建了Forensic Knowledge Database (FKD)和Forensic Chain-of-Thought Dataset (F-CoT)
- 采用三阶段训练方法 (Alignment->SFT->GRPO)
方法论
利用RAG从FKD中检索相关知识,并通过RL在F-CoT上训练MLLM,提升其批判性推理能力,最终实现Deepfake检测。
原文摘要
In Deepfake Detection (DFD) tasks, researchers proposed two types of MLLM-based methods: complementary combination with small DFD detectors, or static forgery knowledge injection.The lack of professional forgery knowledge hinders the performance of these DFD-MLLMs.To solve this, we deeply considered two insightful issues: How to provide high-quality associated forgery knowledge for MLLMs? AND How to endow MLLMs with critical reasoning abilities given noisy reference information? Notably, we attempted to address above two questions with preliminary answers by leveraging the combination of Retrieval-Augmented Generation (RAG) and Reinforcement Learning (RL).Through RAG and RL techniques, we propose the VRAG-DFD framework with accurate dynamic forgery knowledge retrieval and powerful critical reasoning capabilities.Specifically, in terms of data, we constructed two datasets with RAG: Forensic Knowledge Database (FKD) for DFD knowledge annotation, and Forensic Chain-of-Thought Dataset (F-CoT), for critical CoT construction.In terms of model training, we adopt a three-stage training method (Alignment->SFT->GRPO) to gradually cultivate the critical reasoning ability of the MLLM.In terms of performance, VRAG-DFD achieved SOTA and competitive performance on DFD generalization testing.