Multimodal Learning 相关度: 8/10

DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing

Ke Li, Maoliang Li, Jialiang Chen, Jiayu Chen, Zihao Zheng, Shaoqi Wang, Xiang Chen
arXiv: 2604.04875v1 发布: 2026-04-06 更新: 2026-04-06

AI 摘要

DIRECT框架通过层级多智能体规划实现高质量视频混剪,优化视觉和听觉连贯性。

主要贡献

  • 提出多模态连贯性满足问题(MMCSP),并设计DIRECT框架。
  • 构建层级多智能体框架,包含Screenwriter, Director, Editor三个层级。
  • 引入Mashup-Bench基准,包含视觉连贯性和听觉对齐的指标。

方法论

将视频混剪分解为MMCSP,通过层级多智能体架构,逐层优化视频的结构、意图和编辑细节。

原文摘要

Video mashup creation represents a complex video editing paradigm that recomposes existing footage to craft engaging audio-visual experiences, demanding intricate orchestration across semantic, visual, and auditory dimensions and multiple levels. However, existing automated editing frameworks often overlook the cross-level multimodal orchestration to achieve professional-grade fluidity, resulting in disjointed sequences with abrupt visual transitions and musical misalignment. To address this, we formulate video mashup creation as a Multimodal Coherency Satisfaction Problem (MMCSP) and propose the DIRECT framework. Simulating a professional production pipeline, our hierarchical multi-agent framework decomposes the challenge into three cascade levels: the Screenwriter for source-aware global structural anchoring, the Director for instantiating adaptive editing intent and guidance, and the Editor for intent-guided shot sequence editing with fine-grained optimization. We further introduce Mashup-Bench, a comprehensive benchmark with tailored metrics for visual continuity and auditory alignment. Extensive experiments demonstrate that DIRECT significantly outperforms state-of-the-art baselines in both objective metrics and human subjective evaluation. Project page and code: https://github.com/AK-DREAM/DIRECT

标签

video editing multi-agent multimodal

arXiv 分类

cs.CV cs.AI cs.MM