Multimodal Learning 相关度: 9/10

Training-Free Semantic Multi-Object Tracking with Vision-Language Models

Laurence Bonat, Francesco Tonini, Elisa Ricci, Lorenzo Vaquero
arXiv: 2604.14074v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

提出TF-SMOT,一个无需训练的语义多目标跟踪流程,提升视频理解能力。

主要贡献

  • 提出无需训练的SMOT流程TF-SMOT
  • 结合D-FINE, SAM2, InternVideo2.5等预训练模型
  • 在BenSMOT数据集上取得SOTA跟踪性能和更好的视频摘要/字幕质量

方法论

利用预训练模型进行检测、跟踪和视频语言生成,通过LLM进行语义消歧,实现端到端的语义多目标跟踪。

原文摘要

Semantic Multi-Object Tracking (SMOT) extends multi-object tracking with semantic outputs such as video summaries, instance-level captions, and interaction labels, aiming to move from trajectories to human-interpretable descriptions of dynamic scenes. Existing SMOT systems are trained end-to-end, coupling progress to expensive supervision, limiting the ability to rapidly adapt to new foundation models and new interactions. We propose TF-SMOT, a training-free SMOT pipeline that composes pretrained components for detection, mask-based tracking, and video-language generation. TF-SMOT combines D-FINE and the promptable SAM2 segmentation tracker to produce temporally consistent tracklets, uses contour grounding to generate video summaries and instance captions with InternVideo2.5, and aligns extracted interaction predicates to BenSMOT WordNet synsets via gloss-based semantic retrieval with LLM disambiguation. On BenSMOT, TF-SMOT achieves state-of-the-art tracking performance within the SMOT setting and improves summary and caption quality compared to prior art. Interaction recognition, however, remains challenging under strict exact-match evaluation on the fine-grained and long-tailed WordNet label space; our analysis and ablations indicate that semantic overlap and label granularity substantially affect measured performance.

标签

语义多目标跟踪 视频理解 预训练模型 视觉语言模型 无训练

arXiv 分类

cs.CV