When Text Matters: Design Principles for Visual Token Pruning in Vision-Language Model

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant performance degradation in existing visual token pruning methods caused by early text guidance, which restricts the identification of answer-relevant regions. To overcome this limitation, we propose a training-free two-stage pruning strategy that innovatively decouples vision-guided pruning from delayed text-guided re-selection. Specifically, the method first performs preliminary pruning using visual encoder attention, and subsequently refines the final token set via text-to-vision attention at intermediate decoder layers. Extensive experiments across three models and eight benchmarks demonstrate that our approach recovers an average of 11.10 and 16.84 percentage points in performance at 80% and 90% pruning ratios, respectively, while effectively reducing inference latency.
📝 Abstract
Visual token pruning has been widely studied as a practical approach to reducing the computational cost of large vision-language models. However, it struggles to preserve essential visual information, which can lead to substantial performance degradation. In particular, image-based token selection can overlook task-relevant details, while text-guided token selection may fail to capture the text--visual relationships needed for complex reasoning. We find that applying textual guidance too early can limit its ability to identify answer-relevant visual regions, whereas text-to-visual attention becomes more informative at intermediate decoder depths. This finding motivates our training-free method, which separates early vision-guided pruning from deferred text-guided reselection. We first prune visual tokens using vision-encoder attention, retain additional candidates until the decoder midpoint, and then use text-to-visual attention to determine the final visual-token set. Across eight benchmarks and three models, our method outperforms the best-performing baselines by an average of 11.10 and 16.84 percentage points in performance recovery at 80% and 90% pruning, respectively, with comparable or lower LLM-prefill latency than most baselines. The source code is publicly available at https://github.com/kmc3661/DeFT
Problem

Research questions and friction points this paper is trying to address.

Visual Token Pruning
Vision-Language Model
Computational Cost
Text-guided Selection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Token Pruning
Vision-Language Model
Training-free
Text-guided Reselection
Deferred Attention
M
Minchan Kang
Korea Advanced Institute of Science and Technology (KAIST)
K
Kyeonghye Park
Korea Advanced Institute of Science and Technology (KAIST)
S
Seoyoung Cho
Korea Advanced Institute of Science and Technology (KAIST)
D
Daeshik Kim
Korea Advanced Institute of Science and Technology (KAIST)
Y
Yucheol Cho
Hanbat National University