ACPruner: Visual Token Pruning as Biased Attention Coverage Maximization in LVLMs

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the computational efficiency bottleneck in large vision-language models caused by visual token redundancy by proposing ACPruner, a training-free pruning framework. For the first time, this work reformulates token pruning from a global coverage perspective, modeling it as a biased attention coverage maximization problem. Specifically, ACPruner constructs a hybrid importance metric by integrating intra-modal saliency with inter-modal correlation, and derives coverage scores by leveraging the attention patterns of the vision encoder to perform greedy selection for preserving critical information. Extensive experiments on mainstream architectures, including LLaVA and Qwen, demonstrate that the proposed method achieves substantial end-to-end inference acceleration without requiring any additional training, while maintaining superior task performance.
📝 Abstract
Large Vision-Language Models (LVLMs) face significant computational inefficiencies caused by the large number of visual tokens. Existing visual token pruning methods mainly focus on either retaining individually important tokens or selecting mutually diverse ones. In this work, we revisit visual token pruning from a coverage perspective and formulate it as a biased attention coverage maximization problem. The key idea is to select a compact token subset whose encoder-side outgoing attention can jointly cover the image while assigning higher coverage priority to more informative regions. From this perspective, we propose ACPruner, a training-free visual token pruning framework for efficient LVLM inference. ACPruner first estimates token importance by combining intra-modal saliency and inter-modal relevance, then derives token-wise coverage from attention patterns within the vision encoder, and finally performs greedy selection to maximize the proposed coverage objective. Extensive experiments across multiple LVLM backbones, including LLaVA-1.5-7B/13B, LLaVA-NeXT-7B/13B, Qwen2.5-VL-7B, and LLaVA-OneVision-7B, show that ACPruner consistently achieves strong performance retention while delivering substantial end-to-end inference speedups.
Problem

Research questions and friction points this paper is trying to address.

Large Vision-Language Models
Visual Token Pruning
Computational Efficiency
Attention Coverage
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Token Pruning
Attention Coverage Maximization
Training-free
Large Vision-Language Models
Inference Acceleration
X
Xu Li
College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China
Yuxuan Liang
Yuxuan Liang
Assistant Professor, Hong Kong University of Science and Technology (Guangzhou)
Spatio-Temporal Data MiningUrban ComputingUrban AIFoundation ModelsTime Series
Y
Yi Zheng
College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China
Z
Zhe Liu
College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China
X
Xiaolei Chen
College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China
Haotian Chen
Haotian Chen
University of California, Los Angeles
Political EconomyNon-market StrategyAmerican Politics
R
Rui Zhu
College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China
Fan Shi
Fan Shi
Fudan University
Machine LearningDeep Learning
Xiangyang Xue
Xiangyang Xue
Professor of Computer Science, Fudan University
Computer VisionPattern RecognitionMachine Learning