Multimodal Learning 相关度: 9/10

Is CLIP Cross-Eyed? Revealing and Mitigating Center Bias in the CLIP Family

Oscar Chew, Hsiao-Ying Huang, Kunal Jain, Tai-I Chen, Khoa D Doan, Kuan-Hao Huang
arXiv: 2604.05971v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

CLIP模型存在中心偏见,忽略图像边缘物体。通过分析和干预注意力机制,可以缓解这一问题。

主要贡献

  • 揭示CLIP模型家族的中心偏见问题
  • 分析中心偏见的根本原因(表示和注意力)
  • 提出无需训练的策略缓解中心偏见

方法论

通过嵌入分解和注意力图分析,发现信息聚合过程中的信息损失是中心偏见的根源。使用视觉提示和注意力重定向策略缓解。

原文摘要

Recent research has shown that contrastive vision-language models such as CLIP often lack fine-grained understanding of visual content. While a growing body of work has sought to address this limitation, we identify a distinct failure mode in the CLIP family, which we term center bias, that persists even in recent model variants. Specifically, CLIP tends to disproportionately focus on the central region of an image, overlooking important objects located near the boundaries. This limitation is fundamental as failure to recognize relevant objects makes it difficult to perform any sophisticated tasks that depend on those objects. To understand the underlying causes of the limitation, we conduct analyses from both representation and attention perspectives. Using interpretability methods, i.e., embedding decomposition and attention map analysis, we find that relevant concepts especially those associated with off-center objects vanish from the model's embedding in the final representation due to information loss during the aggregation of visual embeddings, particularly the reliance on pooling mechanisms. Finally, we show that this bias can be alleviated with training-free strategies such as visual prompting and attention redistribution by redirecting models' attention to off-center regions.

标签

CLIP 视觉语言模型 中心偏见 注意力机制 可解释性

arXiv 分类

cs.CV cs.CL