🤖 AI Summary
This study addresses the prohibitive inference overhead of vision-language models caused by processing all visual tokens across every layer. We propose P2P, a training-free framework that, for the first time, repurposes activation patching from mechanistic interpretability as an inference acceleration technique. Through forward and backward layer scanning guided by verification, P2P identifies decoder regions where activations can be safely replaced with neutral proxy vectors, achieving computational pruning without token dropping or weight modification. Experiments demonstrate that under a 3% error tolerance, P2P retains 94% of baseline accuracy while reducing floating-point operations by 55%. These findings reveal the non-uniform distribution of visual processing along network depth, offering a principled approach to efficient inference in multimodal architectures.
📝 Abstract
Vision language models (VLMs) incur substantial inference cost because every visual token is processed by the attention and MLP projections of every decoder layer, even when token-specific visual computation is unnecessary at many depths. We introduce Patch-to-Prune (P2P), inspired by Mechanistic Interpretability, a training-free framework that converts activation patching from a diagnostic tool into an inference-time computation bypass. P2P performs validation-guided forward and backward layer sweeps to identify decoder regions whose visual-token projection outputs can be replaced by fixed neutral proxy activation vectors within a user-specified accuracy tolerance. Unlike conventional token-pruning methods, P2P preserves the sequence length, token order, positional information, attention mask, and residual pathways, thereby pruning computation without removing tokens or modifying the pretrained model weights. We evaluate P2P on four VLMs from the Qwen2.5-VL and LLaVA families across seven multi-modal benchmarks using mutually disjoint calibration, validation, and test partitions. P2P at a 3% tolerance retains around 94% of dense accuracy while reducing FLOPs by 55%. Beyond these efficiency gains, our layer-wise analysis suggests that visual processing in VLMs is non-uniformly distributed across decoder depth: early and late layers often require little token-specific visual computation, whereas intermediate layers appear to perform most task-relevant visual integration, enabling later reasoning to rely largely on visual information already embedded in shared residual and textual representations. This makes P2P both an efficient inference framework and a causal lens into visual information processing in VLMs.