Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging
AI 摘要
针对视频语言模型时序推理能力下降问题,提出MERIT方法进行层级选择性模型融合,有效恢复时序推理能力。
主要贡献
- 提出MERIT模型融合框架,无需重新训练即可恢复VLM的时序推理能力
- 通过层级选择性融合,在提高时序推理的同时,保持甚至提高时间感知能力
- 在多个视频基准测试上验证了MERIT的有效性和泛化性
方法论
MERIT通过搜索VLM和纯文本LLM之间层级自注意力融合方案,优化时序推理并惩罚时间感知退化。
原文摘要
Multimodal adaptation equips large language models (LLMs) with perceptual capabilities, but often weakens the reasoning ability inherited from language-only pretraining. This trade-off is especially pronounced in video-language models (VLMs), where visual alignment can impair temporal reasoning (TR) over sequential events. We propose MERIT, a training-free, task-driven model merging framework for restoring TR in VLMs. MERIT searches over layer-wise self-attention merging recipes between a VLM and its paired text-only backbone using an objective that improves TR while penalizing degradation in temporal perception (TP). Across three representative VLMs and multiple challenging video benchmarks, MERIT consistently improves TR, preserves or improves TP, and generalizes beyond the search set to four distinct benchmarks. It also outperforms uniform full-model merging and random layer selection, showing that effective recovery depends on selecting the right layers. Interventional masking and frame-level attribution further show that the selected layers are disproportionately important for reasoning and shift model decisions toward temporally and causally relevant evidence. These results show that targeted, perception-aware model merging can effectively restore TR in VLMs without retraining.