🤖 AI Summary
This study addresses the prohibitive inference overhead caused by visual token redundancy in large video models by proposing a training-free dynamic token compression framework. The method formulates token selection as a progressive evidence accumulation process, overcoming the limitations of conventional independent scoring that neglects cross-frame complementarity. Specifically, it allocates a global budget based on temporal novelty and achieves efficient filtering by jointly leveraging local representativeness and subspace complementarity. Experimental results demonstrate that retaining only 25% of tokens preserves 98.6% of the performance on Qwen3-VL while reducing inference latency by 44.7%. Notably, on LLaVA-OV, the proposed approach even surpasses the accuracy of the original uncompressed model.
📝 Abstract
Video Large Language Models (VideoLLMs) have achieved strong video understanding capabilities but incur substantial inference overhead due to the large number of visual tokens. Existing VideoLLM token compression methods largely rely on selection-independent scoring, overlooking cross-frame complementarity and consequently retaining redundant evidence across frames. Instead, we view video token selection as a progressive evidence accumulation process. It aims to retain visual evidence that is individually informative and collectively complementary under a limited token budget. Building on this insight, we introduce GleanVID, a training-free inference acceleration framework for VideoLLMs. Specifically, GleanVID first allocates the global token budget across frames according to temporal novelty and then selects tokens by jointly considering local representativeness and subspace complementarity, thereby preserving richer and less redundant visual evidence. Extensive experiments across diverse VideoLLMs and benchmarks demonstrate that GleanVID consistently achieves state-of-the-art performance. Notably, with only 25% of visual tokens, GleanVID preserves 98.6% of Qwen3-VL's original performance while reducing its prefill latency by 44.7%. On LLaVA-OV-7B, GleanVID at a 25% retention ratio even slightly surpasses the original model.