Multimodal Learning 相关度: 7/10

SurFITR: A Dataset for Surveillance Image Forgery Detection and Localisation

Qizhou Wang, Guansong Pang, Christopher Leckie
arXiv: 2604.07101v1 发布: 2026-04-08 更新: 2026-04-08

AI 摘要

提出了SurFITR数据集,用于提升监控图像伪造检测和定位的性能。

主要贡献

  • 构建了大规模监控图像伪造数据集SurFITR
  • 利用多模态LLM生成逼真的篡改图像
  • 验证了现有检测器在SurFITR上的性能下降
  • 证明了在SurFITR上训练可以显著提高检测性能

方法论

利用多模态LLM驱动的流程生成包含语义信息和细粒度编辑的篡改图像,构建大规模数据集。

原文摘要

We present the Surveillance Forgery Image Test Range (SurFITR), a dataset for surveillance-style image forgery detection and localisation, in response to recent advances in open-access image generation models that raise concerns about falsifying visual evidence. Existing forgery models, trained on datasets with full-image synthesis or large manipulated regions in object-centric images, struggle to generalise to surveillance scenarios. This is because tampering in surveillance imagery is typically localised and subtle, occurring in scenes with varied viewpoints, small or occluded subjects, and lower visual quality. To address this gap, SurFITR provides a large collection of forensically valuable imagery generated via a multimodal LLM-powered pipeline, enabling semantically aware, fine-grained editing across diverse surveillance scenes. It contains over 137k tampered images with varying resolutions and edit types, generated using multiple image editing models. Extensive experiments show that existing detectors degrade significantly on SurFITR, while training on SurFITR yields substantial improvements in both in-domain and cross-domain performance. SurFITR is publicly available on GitHub.

标签

图像伪造检测 监控视频分析 多模态学习 数据集

arXiv 分类

cs.CV cs.AI cs.MM eess.IV