Multimodal Learning 相关度: 9/10

WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval

Yizhuo Xu, Chaojian Yu, Yuanjie Shao, Tongliang Liu, Qinmu Peng, Xinge You
arXiv: 2604.05583v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

针对Composed Image Retrieval过拟合问题,提出权重正则化微调网络WRF4CIR,提升泛化能力。

主要贡献

  • 揭示了VLP-based CIR中存在的显著泛化差距
  • 提出了基于对抗扰动的权重正则化微调方法WRF4CIR
  • 实验证明WRF4CIR能有效缩小泛化差距并提升检索效果

方法论

通过在模型权重上添加对抗扰动进行正则化,增加模型拟合训练数据的难度,从而缓解过拟合。

原文摘要

Composed Image Retrieval (CIR) task aims to retrieve target images based on reference images and modification texts. Current CIR methods primarily rely on fine-tuning vision-language pre-trained models. However, we find that these approaches commonly suffer from severe overfitting, posing challenges for CIR with limited triplet data. To better understand this issue, we present a systematic study of overfitting in VLP-based CIR, revealing a significant and previously overlooked generalization gap across different models and datasets. Motivated by these findings, we introduce WRF4CIR, a Weight-Regularized Fine-tuning network for CIR. Specifically, during the fine-tuning process, we apply adversarial perturbations to the model weights for regularization, where these perturbations are generated in the opposite direction of gradient descent. Intuitively, WRF4CIR increases the difficulty of fitting the training data, which helps mitigate overfitting in CIR under limited triplet supervision. Extensive experiments on benchmark datasets demonstrate that WRF4CIR significantly narrows the generalization gap and achieves substantial improvements over existing methods.

标签

Composed Image Retrieval Overfitting Weight Regularization Fine-tuning Vision-Language Pre-trained Models

arXiv 分类

cs.CV