FlashGaze: Training-Free Multi-Scale Patch Pruning For Efficient Video Understanding

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive computational overhead incurred by spatiotemporal redundancy when multimodal large language models process long videos. We propose a training-free video understanding acceleration method that leverages pixel-space differences as a proxy metric, eliminating the need for auxiliary networks. Prior to Vision Transformer (ViT) encoding, our approach employs quadtree dynamic programming to jointly optimize the dropping, merging, and retention strategies for multi-scale patches, thereby efficiently pruning redundant visual tokens. Experimental results demonstrate that the proposed method preserves 98% of the original accuracy while achieving a 5.4× speedup in ViT encoding and a 17× acceleration during prefilling. Furthermore, GPU memory consumption is reduced by 1.8×. These improvements significantly enhance the efficiency of long video understanding without compromising model performance.
📝 Abstract
Multimodal Large Language Models (MLLMs) have demonstrated strong performance in video understanding, yet efficiently processing long, high-resolution videos remains challenging. Such videos often contain substantial spatiotemporal redundancy, and processing redundant visual tokens can incur avoidable computational overhead. Many existing methods prune visual tokens during or after vision transformer (ViT) encoding, leaving much of the encoding cost unaddressed. Some approaches prune patches before encoding but rely on learned auxiliary networks for patch selection, incurring additional training and inference overhead. To address these limitations, we propose FlashGaze, a training-free method that reduces spatiotemporal redundancy before ViT encoding without introducing auxiliary networks. FlashGaze uses pixel-space differences as a proxy for information loss and employs Quadtree Dynamic Programming to jointly optimize patch dropping, merging, and keeping under a fixed budget. Experiments on two MLLM backbones across multiple benchmarks demonstrate substantial efficiency gains while largely preserving accuracy. On Qwen3-VL-8B, FlashGaze retains 98% of the full-input baseline accuracy on LongVideoBench while achieving up to 5.4x and 17x speedups in ViT encoding and MLLM prefill, respectively, and reducing peak GPU memory usage by a factor of 1.8. These efficiency gains enable the model to process videos with more frames and higher resolutions on the same GPU hardware, unlocking video understanding at scales previously out of reach.
Problem

Research questions and friction points this paper is trying to address.

Video Understanding
Multimodal Large Language Models
Visual Token Pruning
Spatiotemporal Redundancy
Computational Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-Free
Multi-Scale Patch Pruning
Quadtree Dynamic Programming
Video Understanding
Spatiotemporal Redundancy
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Ziye Zhu
Peking University
Y
Yanghao Zhou
Canva Research
L
Lixing Tan
Canva Research
J
Jialiang Kang
Canva Research
S
Shuxuan Li
Canva Research
X
Xiao Yang
Canva Research