Multimodal Learning 相关度: 9/10

CLASP: Class-Adaptive Layer Fusion and Dual-Stage Pruning for Multimodal Large Language Models

Yunkai Dang, Yizhu Jiang, Yifan Jiang, Qi Fan, Yinghuan Shi, Wenbin Li, Yang Gao
arXiv: 2604.12767v1 发布: 2026-04-14 更新: 2026-04-14

AI 摘要

CLASP通过类自适应层融合和双阶段剪枝,有效减少多模态大语言模型中的视觉token冗余。

主要贡献

  • 提出了一种基于类自适应层融合的视觉特征构建方法
  • 设计了一种双阶段剪枝策略,平衡了相关性和覆盖率
  • 实验证明CLASP在多种基准测试中优于现有方法

方法论

CLASP首先融合多层视觉特征,然后进行双阶段剪枝,以类自适应的方式分配token预算,从而实现视觉token的有效缩减。

原文摘要

Multimodal Large Language Models (MLLMs) suffer from substantial computational overhead due to the high redundancy in visual token sequences. Existing approaches typically address this issue using single-layer Vision Transformer (ViT) features and static pruning strategies. However, such fixed configurations are often brittle under diverse instructions. To overcome these limitations, we propose CLASP, a plug-and-play token reduction framework based on class-adaptive layer fusion and dual-stage pruning. Specifically, CLASP first constructs category-specific visual representations through multi-layer vision feature fusion. It then performs dual-stage pruning, allocating the token budget between attention-salient pivot tokens for relevance and redundancy-aware completion tokens for coverage. Through class-adaptive pruning, CLASP enables prompt-conditioned feature fusion and budget allocation, allowing aggressive yet robust visual token reduction. Extensive experiments demonstrate that CLASP consistently outperforms existing methods across a wide range of benchmarks, pruning ratios, and MLLM architectures. Code will be available at https://github.com/Yunkaidang/CLASP.

标签

多模态学习 视觉token剪枝 模型优化

arXiv 分类

cs.CV cs.AI