Multimodal Learning 相关度: 9/10

ID-Selection: Importance-Diversity Based Visual Token Selection for Efficient LVLM Inference

Zhaohong Huang, Wenjing Liu, Yuxin Zhang, Fei Chao, Rongrong Ji
arXiv: 2604.05601v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

提出了一种基于重要性和多样性的视觉token选择方法ID-Selection,用于加速LVLM推理。

主要贡献

  • 提出ID-Selection算法,平衡token重要性和多样性
  • 无需额外训练,在极端剪枝比例下表现优异
  • 在多个LVLM模型和基准测试上验证了有效性

方法论

ID-Selection首先评估token重要性,然后迭代选择高分token,同时抑制相似token的分数,降低冗余。

原文摘要

Recent advances have explored visual token pruning to accelerate the inference of large vision-language models (LVLMs). However, existing methods often struggle to balance token importance and diversity: importance-based methods tend to retain redundant tokens, whereas diversity-based methods may overlook informative ones. This trade-off becomes especially problematic under high reduction ratios, where preserving only a small subset of visual tokens is critical. To address this issue, we propose ID-Selection, a simple yet effective token selection strategy for efficient LVLM inference. The key idea is to couple importance estimation with diversity-aware iterative selection: each token is first assigned an importance score, after which high-scoring tokens are selected one by one while the scores of similar tokens are progressively suppressed. In this way, ID-Selection preserves informative tokens while reducing redundancy in a unified selection process. Extensive experiments across 5 LVLM backbones and 16 main benchmarks demonstrate that ID-Selection consistently achieves superior performance and efficiency, especially under extreme pruning ratios. For example, on LLaVA-1.5-7B, ID-Selection prunes 97.2% of visual tokens, retaining only 16 tokens, while reducing inference FLOPs by over 97% and preserving 91.8% of the original performance, all without additional training.

标签

视觉token选择 LVLM 模型加速 重要性采样 多样性采样

arXiv 分类

cs.CV