Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding
AI 摘要
提出Multi-Stream Scene Script范式,解耦视频信息并显式关联,提升视频理解和生成效果。
主要贡献
- 提出Multi-Stream Scene Script (MTSS) 范式
- 通过Stream Factorization解耦视频信息
- 通过Relational Grounding显式关联解耦信息
- 提升视频理解任务的性能
- 提升多镜头视频生成任务的性能
方法论
将视频分解为参考、镜头、事件和全局流,通过身份和时间链接关联这些流,构建显式结构化场景描述。
原文摘要
Advances in Multimodal Large Language Models (MLLMs) are transforming video captioning from a descriptive endpoint into a semantic interface for both video understanding and generation. However, the dominant paradigm still casts videos as monolithic narrative paragraphs that entangle visual, auditory, and identity information. This dense coupling not only compromises representational fidelity but also limits scalability, since even local edits can trigger global rewrites. To address this structural bottleneck, we propose Multi-Stream Scene Script (MTSS), a novel paradigm that replaces monolithic text with factorized and explicitly grounded scene descriptions. MTSS is built on two core principles: Stream Factorization, which decouples a video into complementary streams (Reference, Shot, Event, and Global), and Relational Grounding, which reconnects these isolated streams through explicit identity and temporal links to maintain holistic video consistency. Extensive experiments demonstrate that MTSS consistently enhances video understanding across various models, achieving an average reduction of 25% in the total error rate on Video-SALMONN-2 and an average performance gain of 67% on the Daily-Omni reasoning benchmark. It also narrows the performance gap between smaller and larger MLLMs, indicating a substantially more learnable caption interface. Finally, even without architectural adaptation, replacing monolithic prompts with MTSS in multi-shot video generation yields substantial human-rated improvements: a 45% boost in cross-shot identity consistency, a 56% boost in audio-visual alignment, and a 71% boost in temporal controllability.