Multimodal Learning 相关度: 9/10

Empowering Video Translation using Multimodal Large Language Models

Bingzheng QU, Kehai Chen, Xuefeng Bai, Min Zhang
arXiv: 2604.11283v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

该论文全面回顾了基于多模态大语言模型(MLLMs)的视频翻译技术,并提出了三角色分类法。

主要贡献

  • 首次系统性地综述了MLLMs在视频翻译中的应用
  • 提出了MLLMs在视频翻译中的三角色分类框架
  • 讨论了视频翻译中MLLMs面临的挑战和未来方向

方法论

该论文通过文献调研和分析,对MLLMs在视频翻译中的应用进行了梳理和总结,并提出了相应的分类框架。

原文摘要

Recent developments in video translation have further enhanced cross-lingual access to video content, with multimodal large language models (MLLMs) playing an increasingly important supporting role. With strong multimodal understanding, reasoning, and generation capabilities, MLLMs-based video translation systems are overcoming the limitations of traditional cascaded pipelines that separately handle automatic speech recognition, machine translation, text-to-speech and lip synchronization. These MLLM-powered approaches not only achieve competitive or superior translation quality, but also demonstrate stronger robustness in zero-shot settings and multi-speaker scenarios, while jointly modeling semantic fidelity, timing, speaker identity, and emotional consistency. However, despite the rapid progress of MLLMs and extensive surveys on general video-language understanding, a focused and systematic review of how MLLMs empower video translation tasks is still lacking. To fill this gap, we provide the first comprehensive overview of MLLMs-based video translation, organized around a three-role taxonomy: 1) Semantic Reasoner, which characterizes how MLLMs perform video understanding, temporal reasoning, and multimodal fusion; 2) Expressive Performer, which analyzes LLM-driven and LLM-augmented techniques for expressive, controllable speech generation; and 3) Visual Synthesizer, which examines different types of video generators for high-fidelity lip-sync and visual alignment. Finally, we discuss open challenges in video understanding, temporal modeling, and multimodal alignment, and outline promising future research directions for MLLMs-powered video translation.

标签

MLLM 视频翻译 多模态学习 语音合成 唇形同步

arXiv 分类

cs.CV