Multimodal Learning 相关度: 9/10

SLQ: Bridging Modalities via Shared Latent Queries for Retrieval with Frozen MLLMs

Haoran Lou, Ziyan Liu, Chunxiao Fan, Yuexin Wu, Yue Ming
arXiv: 2604.13710v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

提出SLQ框架,通过共享潜在查询将冻结的MLLM适配于检索任务,并构建知识推理检索基准KARR-Bench。

主要贡献

  • 提出SLQ框架,高效利用冻结MLLM进行检索
  • 引入共享潜在查询,实现多模态信息融合
  • 构建KARR-Bench基准,评估知识推理检索能力

方法论

使用共享潜在查询作为文本和图像token序列的全局聚合接口,产生统一空间的嵌入,同时保持MLLM主干网络不变。

原文摘要

Multimodal Large Language Models (MLLMs) exhibit strong reasoning and world knowledge, yet adapting them for retrieval remains challenging. Existing approaches rely on invasive parameter updates, such as full fine-tuning and LoRA, which may disrupt the pre-trained semantic space and impair the structured knowledge essential for reasoning. In this work, we argue that adapting MLLMs for retrieval should focus on eliciting pre-trained representations rather than overwriting them. To this end, we propose SLQ, an effective and efficient framework that adapts a frozen MLLM into a retriever through a small set of Shared Latent Queries. Appended to the end of both text and image token sequences, these queries leverage the model's native causal attention to serve as global aggregation interfaces, producing compact embeddings in a unified space while keeping the backbone unchanged. Furthermore, to better evaluate retrieval beyond superficial pattern matching, we construct KARR-Bench, a benchmark designed for knowledge-aware reasoning retrieval. Extensive experiments show that SLQ outperforms full fine-tuning and LoRA on COCO and Flickr30K, while achieving competitive performance on MMEB and yielding substantial gains on KARR-Bench. The results demonstrate that SLQ, which preserves pre-trained representations, provides an effective and efficient framework for adapting MLLMs to retrieval.

标签

MLLM 检索 多模态 知识推理 冻结模型

arXiv 分类

cs.CV