Multimodal Learning 相关度: 9/10

VisPCO: Visual Token Pruning Configuration Optimization via Budget-Aware Pareto-Frontier Learning for Vision-Language Models

Huawei Ji, Yuanhao Sun, Yuan Jin, Cheng Deng, Jiaxin Ding, Luoyi Fu, Xinbing Wang
arXiv: 2604.15188v1 发布: 2026-04-16 更新: 2026-04-16

AI 摘要

VisPCO提出了一种自动优化视觉token剪枝配置的框架,提升视觉-语言模型的效率。

主要贡献

  • 提出了一种Pareto配置优化方法自动搜索最优剪枝配置
  • 利用连续松弛和直通估计器实现基于梯度的搜索
  • 揭示了多步渐进剪枝能够更好地捕捉VLMs的层次压缩结构

方法论

将视觉token剪枝转化为Pareto配置优化问题,通过Augmented Lagrangian方法求解,并使用可学习的核函数探索逐层剪枝模式。

原文摘要

Visual token pruning methods effectively mitigate the quadratic computational growth caused by processing high-resolution images and video frames in vision-language models (VLMs). However, existing approaches rely on predefined pruning configurations without determining whether they achieve computation-performance optimality. In this work, we introduce , a novel framework that formulates visual token pruning as a Pareto configuration optimization problem to automatically identify optimal configurations. Our approach employs continuous relaxation and straight-through estimators to enable gradient-based search, solved via the Augmented Lagrangian method. Extensive experiments across 8 visual benchmarks demonstrate that effectively approximates the empirical Pareto frontier obtained through grid search and generalizes well across various pruning methods and VLM architectures. Furthermore, through learnable kernel functions, we investigate layer-wise pruning patterns and reveal that multi-step progressive pruning captures VLMs' hierarchical compression structure, achieving superior accuracy-efficiency trade-offs compared to single-layer approaches.

标签

视觉-语言模型 token剪枝 Pareto优化 模型压缩

arXiv 分类

cs.CV cs.AI