Multimodal Learning 相关度: 9/10

CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding Space

Sohwi Lim, Lee Hyoseok, Jungjoon Park, Tae-Hyun Oh
arXiv: 2604.11539v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

CLAY提出一种文本条件视觉相似度调制方法,实现高效且多条件图像检索。

主要贡献

  • 提出CLAY方法,无需额外训练实现文本条件视觉相似度调制
  • 构建CLAY-EVAL数据集,用于多条件检索评估
  • 实验验证CLAY在检索精度和计算效率上的优势

方法论

利用预训练VLMs的embedding空间,将视觉特征提取与文本条件过程分离,实现文本条件下的相似度计算。

原文摘要

Human perception of visual similarity is inherently adaptive and subjective, depending on the users' interests and focus. However, most image retrieval systems fail to reflect this flexibility, relying on a fixed, monolithic metric that cannot incorporate multiple conditions simultaneously. To address this, we propose CLAY, an adaptive similarity computation method that reframes the embedding space of pretrained Vision-Language Models (VLMs) as a text-conditional similarity space without additional training. This design separates the textual conditioning process and visual feature extraction, allowing highly efficient and multi-conditioned retrieval with fixed visual embeddings. We also construct a synthetic evaluation dataset CLAY-EVAL, for comprehensive assessment under diverse conditioned retrieval settings. Experiments on standard datasets and our proposed dataset show that CLAY achieves high retrieval accuracy and notable computational efficiency compared to previous works.

标签

视觉语言模型 图像检索 条件检索 相似度学习

arXiv 分类

cs.CV cs.AI