Multimodal Learning 相关度: 9/10

UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding

Joungbin An, Agrim Jain, Kristen Grauman
arXiv: 2604.08522v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

提出UniversalVTG,通过跨数据集预训练和Query Unifier,实现轻量级且通用的视频时序定位模型。

主要贡献

  • 提出UniversalVTG模型,用于通用视频时序定位
  • 引入Query Unifier解决跨数据集训练中的负迁移问题
  • 模型轻量化,性能优于或匹配大型MLLM

方法论

采用大规模跨数据集预训练,利用Query Unifier统一查询格式,并设计高效的 grounding head。

原文摘要

Video temporal grounding (VTG) is typically tackled with dataset-specific models that transfer poorly across domains and query styles. Recent efforts to overcome this limitation have adapted large multimodal language models (MLLMs) to VTG, but their high compute cost and limited video context still hinder long-video grounding. We instead scale unified supervision while keeping the model lightweight. We present UniversalVTG, a single VTG model trained with large-scale cross-dataset pretraining. An offline Query Unifier canonicalizes heterogeneous query formats into a shared declarative space, reducing linguistic mismatch and preventing the negative transfer observed under naïve joint training. Combined with an efficient grounding head, UniversalVTG scales to long, untrimmed videos. Across diverse benchmarks-GoalStep-StepGrounding, Ego4D-NLQ, TACoS, Charades-STA, and ActivityNet-Captions-one UniversalVTG checkpoint achieves state-of-the-art performance versus dedicated VTG models. Moreover, despite being $>100\times$ smaller than recent MLLM-based approaches, UniversalVTG matches or exceeds their accuracy on multiple benchmarks, offering a practical alternative to parameter-heavy MLLMs.

标签

视频时序定位 多模态学习 迁移学习 轻量级模型

arXiv 分类

cs.CV