VPRune: Efficient Training-free Pre-LLM Visual Token Pruning

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决视觉-语言模型中激进剪枝导致的性能下降问题,提出VPRune框架,通过多样性选择、相似性引导回收和位置保留恢复方法,在保持准确度的同时减少推理成本。
📝 Abstract
Visual token pruning is a promising approach to reducing the inference cost of large vision-language models (LVLMs), yet aggressive token reduction often causes substantial performance degradation. We identify three key factors behind this degradation: text-guided selection bias, information loss from discarded tokens, and positional distortion caused by sequence compaction. Based on these observations, we propose \textbf{VPRune}, a training-free pre-LLM pruning framework consisting of visual-only diversity selection, similarity-guided token recycling, and position-preserving restoration. Experiments on FastVLM-1.5B across multiple vision-language benchmarks demonstrate that VPRune achieves a favorable accuracy--compression trade-off, with particularly pronounced advantages under aggressive compression. Furthermore, evaluations on edge-device show that VPRune effectively reduces end-to-end inference latency while maintaining superior task performance, demonstrating its practicality for resource-constrained LVLM deployment.
Problem

Research questions and friction points this paper is trying to address.

visual token pruning
large vision-language models
inference cost
performance degradation
aggressive token reduction
Innovation

Methods, ideas, or system contributions that make the work stand out.

visual token pruning
training-free
diversity selection
token recycling
position-preserving
💼 Related Jobs
No related jobs found.
G
Guangchuan Lv
Northeastern University
D
Dianxing Shi
Beihang University
Dingjie Fu
Dingjie Fu
Huazhong University of Science and Technology
Computer Vision