P4Q: Co-designing Token Pruning and Quantization for Vision-Language Model Acceleration

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the substantial deployment overhead of vision-language models and the lack of synergy arising from the independent optimization of token pruning and quantization. To overcome these limitations, this work proposes a joint optimization framework that introduces quantization-aware token selection alongside pruning-aware calibration. Specifically, the method leverages pseudo-quantized feature statistics to identify salient tokens and subsequently performs quantization calibration based on the distribution of retained tokens, thereby achieving deep coupling between the two processes. Experimental results demonstrate that the proposed approach yields an average 2.8× end-to-end speedup on LLaVA-NeXT while surpassing existing methods in accuracy.
📝 Abstract
Vision language models have achieved strong performance across a wide range of multimodal applications, yet their substantial computational and memory costs hinder efficient deployment. Visual token pruning and post-training quantization reduce inference overhead along two complementary dimensions, namely sequence length and numerical precision. Existing workflows typically optimize these techniques independently or apply them sequentially. Their distinct optimization objectives leave critical interactions unaddressed and constrain the achievable compression performance. We revisit these designs and present P4Q, a practical co-design framework that jointly optimizes visual token pruning and low-bit quantization for efficient VLM inference. First, P4Q introduces a quantization-aware visual token selection strategy before the LLM. It applies fake quantization to copies of the features produced by the projector and selects visual tokens using statistics computed from these fake-quantized features, thereby conditioning the selector's feature-based decisions on simulated low-bit perturbations. Second, P4Q introduces a pruning-aware quantization calibration strategy. It uses the same selection strategy as pruning to calibrate the quantized model on the retained-token distribution, thereby aligning the calibration process with the pruned execution path used during deployment. By coupling these two components, P4Q achieves substantial inference speedups while maintaining comparable task performance, resulting in a better efficiency-accuracy trade-off than independently optimized pipelines. For instance, on LLaVA-NeXT, P4Q achieves an average end-to-end inference speedup of 2.8x across eight distinct test sets, while retaining higher accuracy than prior compression and quantization methods.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Token Pruning
Post-Training Quantization
Model Compression
Inference Acceleration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Model
Token Pruning
Post-Training Quantization
Co-design Framework
Inference Acceleration
🔎 Similar Papers
No similar papers found.