From Attenuation to Attention: Variational Information Flow Manipulation for Fine-Grained Visual Perception
AI 摘要
针对MLLM细粒度视觉感知弱点,提出变分信息流框架VIF,提升模型对细微视觉信息的关注。
主要贡献
- 提出了“视觉衰减”现象,解释MLLM在细粒度感知上的不足
- 提出了Variational Information Flow (VIF)框架,用于增强细粒度视觉感知
- VIF框架可作为插件集成到现有MLLM架构中,具有良好的泛化性
方法论
利用条件变分自编码器(CVAE)将视觉显著性建模为潜在分布,并在网络传播过程中控制信息流,从而提高对细粒度信息的关注。
原文摘要
While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in general visual understanding, they frequently falter in fine-grained perception tasks that require identifying tiny objects or discerning subtle visual relationships. We attribute this limitation to Visual Attenuation: a phenomenon where sparse fine-grained visual signals are prematurely suppressed or diluted by dominant textual tokens during network propagation, resulting in a "loss of focus" during the deep-level decision-making process. Existing input-centric solutions fail to fundamentally reverse this intrinsic mechanism of information loss. To address this challenge, we propose the Variational Information Flow (VIF) framework. Adopting a probabilistic perspective, VIF leverages a Conditional Variational Autoencoder (CVAE) to model the visual saliency relevant to the question-answer pair as a latent distribution. As a plug-and-play module, VIF can be integrated into existing architectures. Extensive evaluations across diverse benchmarks, covering General VQA, fine-grained perception, and visual grounding, demonstrate that VIF yields competitive improvements over previous methods, validating its effectiveness in enhancing the fine-grained perception of MLLMs.