PoRe: Position-Reweighted Visual Token Pruning for Vision Language Models

📅 2025-08-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Visual-language models (VLMs) commonly prune visual tokens based on text–vision attention scores; however, this practice suffers from the “recency bias” inherent in sequence models, leading to excessive retention of bottom-region image tokens and thus imbalanced pruning. This bias stems from the coupling between spatial image positions and token ordering in the sequence. To address this, we propose a position-aware attention reweighting mechanism that requires no architectural modification or additional training: it explicitly encodes spatial coordinates and dynamically recalibrates attention scores to suppress position-induced bias. Our method is fully compatible with mainstream token pruning frameworks and is empirically validated across multiple VLMs. It consistently improves post-pruning performance—e.g., boosting VQAv2 accuracy by +1.8%—while incurring negligible computational overhead. The approach offers a simple, general, and plug-and-play optimization for efficient VLM inference.

Technology Category

Computer Vision: Large Vision ModelsMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language Models

Application Category

User Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingSearch and Retrieval-Augmented AI: Large language models for searchWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Vision-Language Models (VLMs) typically process a significantly larger number of visual tokens compared to text tokens due to the inherent redundancy in visual signals. Visual token pruning is a promising direction to reduce the computational cost of VLMs by eliminating redundant visual tokens. The text-visual attention score is a widely adopted criterion for visual token pruning as it reflects the relevance of visual tokens to the text input. However, many sequence models exhibit a recency bias, where tokens appearing later in the sequence exert a disproportionately large influence on the model's output. In VLMs, this bias manifests as inflated attention scores for tokens corresponding to the lower regions of the image, leading to suboptimal pruning that disproportionately retains tokens from the image bottom. In this paper, we present an extremely simple yet effective approach to alleviate the recency bias in visual token pruning. We propose a straightforward reweighting mechanism that adjusts the attention scores of visual tokens according to their spatial positions in the image. Our method, termed Position-reweighted Visual Token Pruning, is a plug-and-play solution that can be seamlessly incorporated into existing visual token pruning frameworks without any changes to the model architecture or extra training. Extensive experiments on LVLMs demonstrate that our method improves the performance of visual token pruning with minimal computational overhead.
Problem

Research questions and friction points this paper is trying to address.

Addresses recency bias in visual token pruning
Reduces redundant visual tokens in VLMs
Improves pruning by reweighting attention scores spatially
Innovation

Methods, ideas, or system contributions that make the work stand out.

Position-based attention score reweighting mechanism
Plug-and-play solution for existing frameworks
Mitigates recency bias in visual token pruning
💼 Related Jobs
No related jobs found.
K
Kai Zhao
Shanghai University
W
Wubang Yuan
Shanghai University
A
Alex Lingyu Hung
University of California, Los Angeles
Dan Zeng
Dan Zeng
Sun Yat-sen University
Biometricscomputer visiondeep learning