Rethinking Visual Token Compression for Video Large Language Models: A Simple Yet Strong Baseline

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inference inefficiency of large video models caused by massive visual tokens generated from long videos. We propose SimpleCluster, a training-free method for efficient visual token compression. By leveraging position-aware cross-frame clustering and feature mean representation, this approach substantially reduces token count while effectively preserving spatiotemporal feature structures. Our findings demonstrate that a straightforward feature distribution preservation strategy outperforms more complex compression paradigms. Extensive experiments show that SimpleCluster surpasses existing methods across four benchmarks and three mainstream models, exhibiting remarkable robustness even at an extremely low retention rate of 1%.
📝 Abstract
Video Large Language Models (Video LLMs) have achieved remarkable progress in video understanding, but their inference efficiency is constrained by the large number of visual tokens produced by long videos. Recent video token compression methods increasingly introduce sophisticated strategies for token selection, pruning, and merging. This raises a fundamental question: how much of compression performance can be obtained by simply preserving the structure encoded in the visual representations? We investigate this question with SimpleCluster, a simple and training-free baseline that performs position-aware cross-frame clustering in the visual feature space and represents each cluster using the mean of its original visual features. Extensive experiments across four video understanding benchmarks and three representative Video LLMs show that SimpleCluster achieves competitive or superior performance over recent compression methods across a wide range of token retention ratios, with particularly strong robustness under extremely low retention rates (e.g., 1%). To understand this behavior, we analyze the feature space preserved by different compression methods in terms of local approximation fidelity and global coverage. The results show that stronger downstream performance is consistently associated with better preservation of the original visual feature distribution, especially its global coverage. These findings highlight feature-space preservation as an important consideration for video token compression under highly constrained token budgets. Our code is available at https://github.com/xiaozhang79/SimpleCluster.
Problem

Research questions and friction points this paper is trying to address.

Video Large Language Models
Visual Token Compression
Inference Efficiency
Feature-space Preservation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Token Compression
Video Large Language Models
Training-free Clustering
Feature-space Preservation
Cross-frame Clustering
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Xiao Zhang
Xiao Zhang
Central South University
Computer Vision
W
Wang Zeng
SenseTime Research and Tetras.AI
S
Sheng Jin
SenseTime Research and Tetras.AI
W
Wentao Liu
SenseTime Research and Tetras.AI
C
Chen Qian
SenseTime Research and Tetras.AI
Shichao Kan
Shichao Kan
Central South University
Large Vision Language ModelDeep Metric LearningImage RetrievalObject Retrieval