π€ AI Summary
This study addresses the high inference latency in multimodal large language models caused byεlong visual token sequences, noting that existing pruning methods struggle to balance saliency with semantic consistency. By analyzing token representation dynamics to uncover saliency patterns, this work proposes MSDG-Prune, a training-free pruning framework. The method introduces a novel update-dynamics-based group weighting strategy that leverages directional similarity clustering and query-aware weighting to precisely retain critical information. Evaluated on LLaVA-NeXT, MSDG-Prune preserves 91.9% of the original performance using only 5.6% of the tokens, achieving a 7.8Γ prefill speedup while demonstrating strong generalization capabilities across diverse tasks.
π Abstract
Multimodal large language models (MLLMs) incur high inference latency from long visual token sequences. Existing pruning methods commonly use attention maps or output features to estimate token importance or redundancy. Several recent approaches also exploit representation changes, but when and how these changes reflect foreground saliency and semantic consistency remain insufficiently understood. We analyze visual token representation dynamics across encoder depth and uncover two findings. First, the relationship between token update magnitudes and foreground saliency is layer-dependent: large token updates concentrate on foreground regions in two depth intervals, separated by several sink-dominated layers at intermediate depths. Second, similarities between token update directions better distinguish same-class from different-class tokens than those between encoder output features. Building on these findings, we propose MSDG-Prune, a training-free method that uses update magnitudes and directions to preserve salient and diverse visual information. Specifically, we group tokens by update-direction similarity and use query-weighted saliency derived from update magnitudes across a chosen depth window for group-wise token pruning. Extensive experiments across four MLLMs demonstrate the effectiveness and generalizability of MSDG-Prune. On LLaVA-NeXT, it retains 91.9% of uncompressed performance on average with only 5.6% of visual tokens, while achieving a 7.8x prefilling speedup. Code is available at https://github.com/liweixuan-hitsz/MSDG-Prune.