Multimodal Learning 相关度: 9/10

Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference

Imanol Miranda, Ander Salaberria, Eneko Agirre, Gorka Azkune
arXiv: 2604.11496v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

通过优化推理阶段的对齐方式,提升双编码器视觉语言模型在组合性任务上的泛化能力。

主要贡献

  • 提出全局嵌入匹配是双编码器VLM的关键瓶颈
  • 引入轻量级Transformer学习局部对齐,无需更新预训练编码器
  • 证明局部对齐能显著提升域外组合性基准测试的表现

方法论

通过控制诊断实验,分析现有模型的不足,并引入Transformer学习图像区域和文本片段的局部对齐。

原文摘要

Dual-encoder Vision-Language Models (VLMs) such as CLIP are often characterized as bag-of-words systems due to their poor performance on compositional benchmarks. We argue that this limitation may stem less from deficient representations than from the standard inference protocol based on global cosine similarity. First, through controlled diagnostic experiments, we show that explicitly enforcing fine-grained region-segment alignment at inference dramatically improves compositional performance without updating pretrained encoders. We then introduce a lightweight transformer that learns such alignments directly from frozen patch and token embeddings. Comparing against full fine-tuning and prior end-to-end compositional training methods, we find that although these approaches improve in-domain retrieval, their gains do not consistently transfer under distribution shift. In contrast, learning localized alignment over frozen representations matches full fine-tuning on in-domain retrieval while yielding substantial improvements on controlled out-of-domain compositional benchmarks. These results identify global embedding matching as a key bottleneck in dual-encoder VLMs and highlight the importance of alignment mechanisms for robust compositional generalization.

标签

视觉语言模型 组合性 双编码器 推理 对齐

arXiv 分类

cs.CV cs.CL cs.LG