Multimodal Learning 相关度: 9/10

Beyond Semantics: Uncovering the Physics of Fakes via Universal Physical Descriptors for Cross-Modal Synthetic Detection

Mei Qiu, Jianqiang Zhao, Yanyun Qu
arXiv: 2604.04608v1 发布: 2026-04-06 更新: 2026-04-06

AI 摘要

该论文提出了一种基于物理特征的跨模态合成检测方法,提升了AIGC图像检测的鲁棒性和准确性。

主要贡献

  • 提出了一种新的基于物理特征的AIGC图像检测方法。
  • 探索并选择了五个核心物理特征用于区分真实图像和AI生成图像。
  • 将物理特征集成到CLIP模型中,提高了检测性能并减轻了语言信息的不确定性。

方法论

通过分析GANs和扩散模型生成的图像,提取物理特征,结合文本信息,使用CLIP进行训练和检测。

原文摘要

The rapid advancement of AI generated content (AIGC) has blurred the boundaries between real and synthetic images, exposing the limitations of existing deepfake detectors that often overfit to specific generative models. This adaptability crisis calls for a fundamental reexamination of the intrinsic physical characteristics that distinguish natural from AI-generated images. In this paper, we address two critical research questions: (1) What physical features can stably and robustly discriminate AI generated images across diverse datasets and generative architectures? (2) Can these objective pixel-level features be integrated into multimodal models like CLIP to enhance detection performance while mitigating the unreliability of language-based information? To answer these questions, we conduct a comprehensive exploration of 15 physical features across more than 20 datasets generated by various GANs and diffusion models. We propose a novel feature selection algorithm that identifies five core physical features including Laplacian variance, Sobel statistics, and residual noise variance that exhibit consistent discriminative power across all tested datasets. These features are then converted into text encoded values and integrated with semantic captions to guide image text representation learning in CLIP. Extensive experiments demonstrate that our method achieves state-of-the-art performance on multiple Genimage benchmarks, with near-perfect accuracy (99.8%) on datasets such as Wukong and SDv1.4. By bridging pixel level authenticity with semantic understanding, this work pioneers the use of physically grounded features for trustworthy vision language modeling and opens new directions for mitigating hallucinations and textual inaccuracies in large multimodal models.

标签

AIGC检测 物理特征 跨模态学习 CLIP

arXiv 分类

cs.CV