Reasoning Dynamics and the Limits of Monitoring Modality Reliance in Vision-Language Models
AI 摘要
研究视觉语言模型在推理过程中的信息整合方式,揭示了其对文本线索的过度依赖及CoT的可解释性局限。
主要贡献
- 揭示了VLMs在推理中存在答案惯性现象
- 评估了推理训练对模型纠正错误答案的影响
- 发现VLMs易受误导性文本线索的影响,即使视觉证据充足
方法论
通过追踪CoT的置信度、测量推理的纠正效果、以及使用误导性文本线索进行控制干预,分析了18个VLMs的推理动态。
原文摘要
Recent advances in vision language models (VLMs) offer reasoning capabilities, yet how these unfold and integrate visual and textual information remains unclear. We analyze reasoning dynamics in 18 VLMs covering instruction-tuned and reasoning-trained models from two different model families. We track confidence over Chain-of-Thought (CoT), measure the corrective effect of reasoning, and evaluate the contribution of intermediate reasoning steps. We find that models are prone to answer inertia, in which early commitments to a prediction are reinforced, rather than revised during reasoning steps. While reasoning-trained models show stronger corrective behavior, their gains depend on modality conditions, from text-dominant to vision-only settings. Using controlled interventions with misleading textual cues, we show that models are consistently influenced by these cues even when visual evidence is sufficient, and assess whether this influence is recoverable from CoT. Although this influence can appear in the CoT, its detectability varies across models and depends on what is being monitored. Reasoning-trained models are more likely to explicitly refer to the cues, but their longer and fluent CoTs can still appear visually grounded while actually following textual cues, obscuring modality reliance. In contrast, instruction-tuned models refer to the cues less explicitly, but their shorter traces reveal inconsistencies with the visual input. Taken together, these findings indicate that CoT provides only a partial view of how different modalities drive VLM decisions, with important implications for the transparency and safety of multimodal systems.