ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling
AI 摘要
ControlFoley提出了一种可控的跨模态视频到音频生成框架,解决了跨模态冲突和风格控制问题。
主要贡献
- 提出统一的V2A框架ControlFoley
- 引入联合视觉编码范式和时域-音色解耦
- 设计了模态鲁棒训练方案和VGGSound-TVC基准
方法论
结合CLIP和音视频编码器进行联合视觉编码,解耦时域和音色信息,并采用模态鲁棒训练方案。
原文摘要
Recent advances in video-to-audio (V2A) generation enable high-quality audio synthesis from visual content, yet achieving robust and fine-grained controllability remains challenging. Existing methods suffer from weak textual controllability under visual-text conflict and imprecise stylistic control due to entangled temporal and timbre information in reference audio. Moreover, the lack of standardized benchmarks limits systematic evaluation. We propose ControlFoley, a unified multimodal V2A framework that enables precise control over video, text, and reference audio. We introduce a joint visual encoding paradigm that integrates CLIP with a spatio-temporal audio-visual encoder to improve alignment and textual controllability. We further propose temporal-timbre decoupling to suppress redundant temporal cues while preserving discriminative timbre features. In addition, we design a modality-robust training scheme with unified multimodal representation alignment (REPA) and random modality dropout. We also present VGGSound-TVC, a benchmark for evaluating textual controllability under varying degrees of visual-text conflict. Extensive experiments demonstrate state-of-the-art performance across multiple V2A tasks, including text-guided, text-controlled, and audio-controlled generation. ControlFoley achieves superior controllability under cross-modal conflict while maintaining strong synchronization and audio quality, and shows competitive or better performance compared to an industrial V2A system. Code, models, datasets, and demos are available at: https://yjx-research.github.io/ControlFoley/.