Multimodal Learning 相关度: 9/10

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models

Ami Baid, Zihui Xue, Kristen Grauman
arXiv: 2604.14129v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

提出音频对比偏好优化ACPO,解决音视频语言模型中视频驱动的音频幻觉问题,提升音频的可靠性。

主要贡献

  • 提出Audio-Contrastive Preference Optimization (ACPO)框架
  • 设计输出对比目标,惩罚伪装成音频事实的视觉描述
  • 设计输入对比目标,惩罚对真实音频信号不变的生成
  • 实验证明ACPO能提高音频接地性,减轻音频幻觉

方法论

通过双轴偏好学习,设计输出和输入对比目标,优化模型生成更符合真实音频信息的描述。

原文摘要

While Audio-Visual Language Models (AVLMs) have achieved remarkable progress over recent years, their reliability is bottlenecked by cross-modal hallucination. A particularly pervasive manifestation is video-driven audio hallucination: models routinely exploit visual shortcuts to hallucinate expected sounds, discarding true auditory evidence. To counteract this deeply ingrained visual dominance, we propose Audio-Contrastive Preference Optimization (ACPO). This dual-axis preference learning framework introduces an output-contrastive objective to penalize visual descriptions masquerading as audio facts, alongside an input-contrastive objective that swaps audio tracks to explicitly penalize generation invariant to the true auditory signal. Extensive experiments demonstrate that ACPO establishes highly faithful audio grounding and mitigates audio hallucination without compromising overarching multimodal capabilities.

标签

音视频语言模型 多模态学习 音频幻觉 对比学习 偏好优化

arXiv 分类

cs.CV