🤖 AI Summary
This study addresses the significant performance degradation in existing visual token pruning methods caused by early text guidance, which restricts the identification of answer-relevant regions. To overcome this limitation, we propose a training-free two-stage pruning strategy that innovatively decouples vision-guided pruning from delayed text-guided re-selection. Specifically, the method first performs preliminary pruning using visual encoder attention, and subsequently refines the final token set via text-to-vision attention at intermediate decoder layers. Extensive experiments across three models and eight benchmarks demonstrate that our approach recovers an average of 11.10 and 16.84 percentage points in performance at 80% and 90% pruning ratios, respectively, while effectively reducing inference latency.
📝 Abstract
Visual token pruning has been widely studied as a practical approach to reducing the computational cost of large vision-language models. However, it struggles to preserve essential visual information, which can lead to substantial performance degradation. In particular, image-based token selection can overlook task-relevant details, while text-guided token selection may fail to capture the text--visual relationships needed for complex reasoning. We find that applying textual guidance too early can limit its ability to identify answer-relevant visual regions, whereas text-to-visual attention becomes more informative at intermediate decoder depths. This finding motivates our training-free method, which separates early vision-guided pruning from deferred text-guided reselection. We first prune visual tokens using vision-encoder attention, retain additional candidates until the decoder midpoint, and then use text-to-visual attention to determine the final visual-token set. Across eight benchmarks and three models, our method outperforms the best-performing baselines by an average of 11.10 and 16.84 percentage points in performance recovery at 80% and 90% pruning, respectively, with comparable or lower LLM-prefill latency than most baselines. The source code is publicly available at https://github.com/kmc3661/DeFT