🤖 AI Summary
To address the high computational cost, memory footprint, and latency of parameter-efficient fine-tuning (PEFT) for Vision Transformers (ViTs) during inference, this paper proposes a semantic-aware sparse token processing and dense cross-layer Adapter co-architecture. Methodologically: (1) tokens are dynamically sparsified layer-wise based on attention weights and gradient-based semantic importance scores, preserving semantically critical tokens while merging redundant ones; (2) dense shallow-to-deep Adapter modules are introduced to integrate multi-level local features across layers. This work is the first to jointly optimize PEFT efficiency in terms of GFLOPs, GPU memory consumption, and inference latency. Evaluated on VTAB-1K and multiple image/video benchmarks, our approach reduces GFLOPs by 30–38% for ViT-B, significantly lowers inference latency, and achieves state-of-the-art accuracy.
📝 Abstract
Parameter-efficient fine-tuning (PEFT) has emerged as a popular solution for adapting pre-trained Vision Transformer (ViT) models to downstream applications. While current PEFT methods have achieved parameter efficiency, they overlook the efficiency of computation and GPU memory during both fine-tuning and inference, falling short of practical requirements. In this paper, we propose extbf{Sparse-Tuning}, a novel PEFT method that accounts for the information redundancy in images and videos to boost the above efficiency. By sparsely preserving the semantic-relevant tokens and merging irrelevant ones, Sparse-Tuning minimizes the quantity of tokens processed at each layer, leading to a quadratic reduction in computational and memory overhead. To align our token sparsification strategy suitably with fine-tuning purposes, we further design Dense Adapters that establish dense connections from shallow layers to deeper layers. These Dense Adapters integrate multi-level local features to enrich the current tokens, improving both token preservation and model adaptation. Empirical results on VTAB-1K, three image datasets, and two video datasets show that our Sparse-Tuning reduces GFLOPs to extbf{62%-70%} of the original ViT-B while achieving state-of-the-art performance. Source code is available at url{https://github.com/liuting20/Sparse-Tuning}.