FlashGaze: Training-Free Multi-Scale Patch Pruning For Efficient Video Understanding
This study addresses the prohibitive computational overhead incurred by spatiotemporal redundancy when multimodal large language models process long videos. We propose a training-free video understanding acceleration method that leverages pixel-space differences as a proxy metric, eliminating the need for auxiliary networks. Prior to Vision Transformer (ViT) encoding, our approach employs quadtree dynamic programming to jointly optimize the dropping, merging, and retention strategies for multi-scale patches, thereby efficiently pruning redundant visual tokens. Experimental results demonstrate that the proposed method preserves 98% of the original accuracy while achieving a 5.4× speedup in ViT encoding and a 17× acceleration during prefilling. Furthermore, GPU memory consumption is reduced by 1.8×. These improvements significantly enhance the efficiency of long video understanding without compromising model performance.