Efficient Token Compression for Vision Transformer with Spatial Information Preserved

📅 2025-03-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the high computational and memory overhead of Vision Transformers—hindering their deployment in resource-constrained scenarios—this paper proposes a hierarchical “prune-merge” token compression framework. Our method introduces three key innovations: (1) a gradient-weighted attention scoring mechanism that dynamically evaluates token importance during training; (2) learnable merge/reconstruction matrices coupled with residual connections, enabling structured reconstruction of pruned tokens; and (3) end-to-end joint optimization guided by global gradient sensitivity, automatically discovering optimal compression architectures. Evaluated on ImageNet-1K, our approach accelerates DeiT-Small inference by 1.64× with only a 0.2% top-1 accuracy drop. On ADE20K semantic segmentation, it significantly outperforms existing token compression methods. The framework achieves efficient yet accurate vision modeling without architectural modification, offering a principled pathway toward lightweight ViT deployment.

Technology Category

Computer Vision: Large Vision ModelsMachine Learning: Learning on the Edge & Model CompressionSearch and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphs
📝 Abstract
Token compression is essential for reducing the computational and memory requirements of transformer models, enabling their deployment in resource-constrained environments. In this work, we propose an efficient and hardware-compatible token compression method called Prune and Merge. Our approach integrates token pruning and merging operations within transformer models to achieve layer-wise token compression. By introducing trainable merge and reconstruct matrices and utilizing shortcut connections, we efficiently merge tokens while preserving important information and enabling the restoration of pruned tokens. Additionally, we introduce a novel gradient-weighted attention scoring mechanism that computes token importance scores during the training phase, eliminating the need for separate computations during inference and enhancing compression efficiency. We also leverage gradient information to capture the global impact of tokens and automatically identify optimal compression structures. Extensive experiments on the ImageNet-1k and ADE20K datasets validate the effectiveness of our approach, achieving significant speed-ups with minimal accuracy degradation compared to state-of-the-art methods. For instance, on DeiT-Small, we achieve a 1.64$ imes$ speed-up with only a 0.2% drop in accuracy on ImageNet-1k. Moreover, by compressing segmenter models and comparing with existing methods, we demonstrate the superior performance of our approach in terms of efficiency and effectiveness. Code and models have been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/prune_and_merge.
Problem

Research questions and friction points this paper is trying to address.

Reduces computational and memory needs for vision transformers
Preserves spatial information during token compression
Enables efficient deployment in resource-limited environments
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prune and Merge token compression method
Gradient-weighted attention scoring mechanism
Trainable merge and reconstruct matrices
J
Junzhu Mao
School of Computer Science and Engineering, Nanjing University of Science and Technology, Nanjing 210094, China
Y
Yang Shen
School of Computer Science and Engineering, Nanjing University of Science and Technology, Nanjing 210094, China
Jinyang Guo
Jinyang Guo
The University of Sydney
Deep LearningEfficient MethodsEdge Computing
Y
Yazhou Yao
School of Computer Science and Engineering, Nanjing University of Science and Technology, Nanjing 210094, China
X
Xiansheng Hua
Terminus Group, Beijing, 100027, China