Multimodal Learning 相关度: 9/10

AIM: Asymmetric Information Masking for Visual Question Answering Continual Learning

Peifeng Zhang, Zice Qiu, Donghua Yu, Shilei Cao, Juepeng Zheng, Yutong Lu, Haohuan Fu
arXiv: 2604.14779v1 发布: 2026-04-16 更新: 2026-04-16

AI 摘要

针对VQA持续学习中视觉语言模型的不对称性,提出非对称信息掩码(AIM)方法,平衡稳定性和可塑性。

主要贡献

  • 提出Asymmetric Information Masking (AIM)方法
  • 针对VQA持续学习中的视觉语言模型不对称性问题
  • 在VQA v2和GQA数据集上取得SOTA性能

方法论

针对不同模态的敏感度,使用目标掩码来平衡稳定性和可塑性,缓解灾难性遗忘问题。

原文摘要

In continual visual question answering (VQA), existing Continual Learning (CL) methods are mostly built for symmetric, unimodal architectures. However, modern Vision-Language Models (VLMs) violate this assumption, as their trainable components are inherently asymmetric. This structural mismatch renders VLMs highly prone to catastrophic forgetting when learning from continuous data streams. Specifically, the asymmetry causes standard global regularization to favor the massive language decoder during optimization, leaving the smaller but critical visual projection layers highly vulnerable to interference. Consequently, this localized degradation leads to a severe loss of compositional reasoning capabilities. To address this, we propose Asymmetric Information Masking (AIM), which balances stability and plasticity by applying targeted masks based on modality-specific sensitivity. Experiments on VQA v2 and GQA under continual VQA settings show that AIM achieves state-of-the-art performance in both Average Performance (AP) and Average Forgetting (AF), while better preserving generalization to novel skill-concept compositions.

标签

VQA Continual Learning Multimodal Learning Asymmetric Information Masking

arXiv 分类

cs.CV cs.CL