Rethinking Visual Token Compression for Video Large Language Models: A Simple Yet Strong Baseline
This study addresses the inference inefficiency of large video models caused by massive visual tokens generated from long videos. We propose SimpleCluster, a training-free method for efficient visual token compression. By leveraging position-aware cross-frame clustering and feature mean representation, this approach substantially reduces token count while effectively preserving spatiotemporal feature structures. Our findings demonstrate that a straightforward feature distribution preservation strategy outperforms more complex compression paradigms. Extensive experiments show that SimpleCluster surpasses existing methods across four benchmarks and three mainstream models, exhibiting remarkable robustness even at an extremely low retention rate of 1%.