Multimodal Learning 相关度: 9/10

Scene Change Detection with Vision-Language Representation Learning

Diwei Sheng, Vijayraj Gohil, Satyam Gaba, Zihan Liu, Giles Hamilton-Fletcher, John-Ross Rizzo, Yongqing Liang, Chen Feng
arXiv: 2604.11402v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

提出LangSCD框架,利用视觉-语言模型进行场景变化检测,并构建了多类别变化标注的数据集NYC-CD。

主要贡献

  • 提出LangSCD框架,融合视觉和语言信息进行场景变化检测
  • 引入几何-语义匹配模块,提升预测mask的语义一致性和空间完整性
  • 构建了包含多类别变化的场景变化检测数据集NYC-CD

方法论

利用视觉-语言模型生成场景变化的文本描述,与视觉特征融合,并通过几何-语义匹配模块优化预测结果。

原文摘要

Scene change detection (SCD) is crucial for urban monitoring and navigation but remains challenging in real-world environments due to lighting variations, seasonal shifts, viewpoint differences, and complex urban layouts. Existing methods rely primarily on low-level visual features, limiting their ability to accurately identify changed objects amid the visual complexity of urban scenes. In this paper, we propose LangSCD, a vision-language framework for scene change detection that overcomes this single-modal limitation by incorporating semantic reasoning through language. Our approach introduces a modular language component that leverages vision-language models (VLMs) to generate textual descriptions of scene changes, which are fused with visual features through a cross-modal feature enhancer. We further introduce a geometric-semantic matching module that refines the predicted masks by enforcing semantic consistency and spatial completeness. Existing real-world scene change detection benchmarks provide only binary change annotations, which are insufficient for downstream applications requiring fine-grained understanding of scene dynamics. To address this limitation, we introduce NYC-CD, a large-scale dataset of 8,122 real-world image pairs collected in New York City with multiclass change annotations generated through a semi-automatic pipeline. Extensive experiments across multiple street-view benchmarks demonstrate that our language and matching modules consistently improve existing change-detection architectures, achieving state-of-the-art performance and highlighting the value of integrating linguistic reasoning with visual representations for robust scene change detection.

标签

场景变化检测 视觉-语言模型 多模态学习 数据集

arXiv 分类

cs.CV