Video-guided Machine Translation with Global Video Context
AI 摘要
提出一种全局视频引导的多模态翻译框架,提升长视频翻译效果。
主要贡献
- 提出全局视频引导的多模态翻译框架
- 使用预训练语义编码器和向量数据库检索字幕
- 设计区域感知跨模态注意力机制
方法论
利用语义编码器检索相关视频片段构建上下文,并通过注意力机制和区域感知跨模态注意力进行翻译。
原文摘要
Video-guided Multimodal Translation (VMT) has advanced significantly in recent years. However, most existing methods rely on locally aligned video segments paired one-to-one with subtitles, limiting their ability to capture global narrative context across multiple segments in long videos. To overcome this limitation, we propose a globally video-guided multimodal translation framework that leverages a pretrained semantic encoder and vector database-based subtitle retrieval to construct a context set of video segments closely related to the target subtitle semantics. An attention mechanism is employed to focus on highly relevant visual content, while preserving the remaining video features to retain broader contextual information. Furthermore, we design a region-aware cross-modal attention mechanism to enhance semantic alignment during translation. Experiments on a large-scale documentary translation dataset demonstrate that our method significantly outperforms baseline models, highlighting its effectiveness in long-video scenarios.