Multimodal Learning 相关度: 9/10

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding

Ashutosh Kumar, Rajat Saini, Jingjing Pan, Mustafa Erdogan, Mingfang Zhang, Betty Le Dem, Norimasa Kobori, Quan Kong
arXiv: 2604.08337v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

InstAP通过实例感知的预训练框架,提升视觉-语言模型在时空理解上的细粒度推理能力。

主要贡献

  • 提出了Instance-Aware Pre-training (InstAP)框架
  • 构建了大规模的 InstVL 数据集
  • 证明了实例感知预训练能提高实例级别和全局理解能力

方法论

InstAP通过联合优化全局视觉-文本对齐和细粒度的实例级别对比对齐,提升模型对时空区域的理解。

原文摘要

Current vision-language pre-training (VLP) paradigms excel at global scene understanding but struggle with instance-level reasoning due to global-only supervision. We introduce InstAP, an Instance-Aware Pre-training framework that jointly optimizes global vision-text alignment and fine-grained, instance-level contrastive alignment by grounding textual mentions to specific spatial-temporal regions. To support this, we present InstVL, a large-scale dataset (2 million images, 50,000 videos) with dual-granularity annotations: holistic scene captions and dense, grounded instance descriptions. On the InstVL benchmark, InstAP substantially outperforms existing VLP models on instance-level retrieval, and also surpasses a strong VLP baseline trained on the exact same data corpus, isolating the benefit of our instance-aware objective. Moreover, instance-centric pre-training improves global understanding: InstAP achieves competitive zero-shot performance on multiple video benchmarks, including MSR-VTT and DiDeMo. Qualitative visualizations further show that InstAP localizes textual mentions to the correct instances, while global-only models exhibit more diffuse, scene-level attention.

标签

视觉语言预训练 时空理解 实例感知 对比学习

arXiv 分类

cs.CV cs.AI