V-CoLA: Vision Token Compression with Linear Attention

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the substantial computational overhead of visual tokens in vision-language models and the incompatibility of existing compression methods with linear attention hybrid architectures. To this end, we propose V-CoLA, a training-free visual token compression framework. V-CoLA introduces a novel uniqueness-aware importance criterion tailored for linear attention, integrated with an adaptive token merging strategy and a chunked-parallel-compatible efficient implementation mechanism to achieve lossless visual token compression. Experimental results demonstrate that retaining only 50% of the visual tokens preserves 99.5% of the original performance, while maintaining over 88% even when compressed to 12.5%. Furthermore, V-CoLA yields a 1.86× to 6.15× acceleration during the prefilling stage, establishing it as a highly effective solution for efficient vision-language modeling.
📝 Abstract
Vision-language models (VLMs) have demonstrated impressive capabilities but suffer from substantial computational overhead, as vision tokens dominate the input sequence. This motivates vision token compression as a key direction to alleviate the burden. However, with the emergence of hybrid architectures incorporating linear attention (\eg, Qwen3.5), prior methods designed for softmax attention struggle to generalize. Our analysis reveals that both attention- and similarity-based approaches suffer notable performance degradation, underscoring the urgent need for compression methods tailored to this regime. To this end, we propose \textbf{V-CoLA}, an efficient training-free token compression framework specifically designed for linear attention. V-CoLA introduces a novel \textit{uniqueness-aware importance criterion} for identifying critical vision tokens, coupled with an \textit{adaptive token merging strategy} that performs compression. All components are optimized at the implementation level to remain compatible with the chunk-wise parallelism of linear attention, ensuring strong practical value. Extensive experiments across multiple benchmarks demonstrate the superiority of V-CoLA: it achieves 99.5\% of the original performance with only 50.0\% of vision tokens, and over 88.0\% with as few as 12.5\%, while delivering a 1.86$\times$ to 6.15$\times$ prefill speedup.
Problem

Research questions and friction points this paper is trying to address.

Vision-language models
Vision token compression
Linear attention
Computational overhead
Hybrid architectures
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision Token Compression
Linear Attention
Training-free
Token Merging
Vision-Language Models
💼 Related Jobs
No related jobs found.
Hao Jiang
Hao Jiang
Alibaba Group
LLM & AIGC
Y
Yiru Mao
Alibaba Cloud Computing, Alibaba Group
T
Tianpeng Bu
Alibaba Cloud Computing, Alibaba Group
Hao Zhou
Hao Zhou
Alibaba Group | PhD, SJTU
computer visionvideo understandingAIGC
Hongtao Duan
Hongtao Duan
Nanjing Institute of Geography and Limnology, Chinese Academy of Sciences
Ocean color remote sensingLake remote sensing
Wang Jing
Wang Jing
Alibaba Cloud Computing, Alibaba Group
B
Bowen Xu
Alibaba Cloud Computing, Alibaba Group
X
Xin Chen
Alibaba Cloud Computing, Alibaba Group
L
Lulu Hu
Alibaba Cloud Computing, Alibaba Group
B
Bin Yang
Alibaba Cloud Computing, Alibaba Group
Y
Yongliang Tao
Alibaba Cloud Computing, Alibaba Group
M
Minying Zhang
Alibaba Cloud Computing, Alibaba Group