Multimodal Learning 相关度: 9/10

Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator

Luozheng Qin, Jia Gong, Qian Qiao, Tianjiao Li, Li Xu, Haoyu Pan, Chao Qu, Zhiyu Tan, Hao Li
arXiv: 2604.08121v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

Uni-ViGU通过扩展视频生成器,统一视频生成和理解任务,实现了有竞争力的性能。

主要贡献

  • 提出了基于扩散模型的统一视频生成和理解框架Uni-ViGU
  • 引入了统一的流方法,同时处理视频和文本的生成
  • 设计了基于MoE的框架,增强文本生成能力并保留生成先验

方法论

通过扩展视频生成器,利用连续和离散流匹配实现多模态生成,并通过双向训练机制将生成知识迁移到理解任务。

原文摘要

Unified multimodal models integrating visual understanding and generation face a fundamental challenge: visual generation incurs substantially higher computational costs than understanding, particularly for video. This imbalance motivates us to invert the conventional paradigm: rather than extending understanding-centric MLLMs to support generation, we propose Uni-ViGU, a framework that unifies video generation and understanding by extending a video generator as the foundation. We introduce a unified flow method that performs continuous flow matching for video and discrete flow matching for text within a single process, enabling coherent multimodal generation. We further propose a modality-driven MoE-based framework that augments Transformer blocks with lightweight layers for text generation while preserving generative priors. To repurpose generation knowledge for understanding, we design a bidirectional training mechanism with two stages: Knowledge Recall reconstructs input prompts to leverage learned text-video correspondences, while Capability Refinement fine-tunes on detailed captions to establish discriminative shared representations. Experiments demonstrate that Uni-ViGU achieves competitive performance on both video generation and understanding, validating generation-centric architectures as a scalable path toward unified multimodal intelligence. Project Page and Code: https://fr0zencrane.github.io/uni-vigu-page/.

标签

视频生成 视频理解 多模态学习 扩散模型 统一框架

arXiv 分类

cs.CV cs.AI