Sparse-Tuning: Adapting Vision Transformers with Efficient Fine-tuning and Inference

📅 2024-05-23
🏛️ arXiv.org
📈 Citations: 16
✨ Influential: 1
📄 PDF
🤖 AI Summary
To address the high computational cost, memory footprint, and latency of parameter-efficient fine-tuning (PEFT) for Vision Transformers (ViTs) during inference, this paper proposes a semantic-aware sparse token processing and dense cross-layer Adapter co-architecture. Methodologically: (1) tokens are dynamically sparsified layer-wise based on attention weights and gradient-based semantic importance scores, preserving semantically critical tokens while merging redundant ones; (2) dense shallow-to-deep Adapter modules are introduced to integrate multi-level local features across layers. This work is the first to jointly optimize PEFT efficiency in terms of GFLOPs, GPU memory consumption, and inference latency. Evaluated on VTAB-1K and multiple image/video benchmarks, our approach reduces GFLOPs by 30–38% for ViT-B, significantly lowers inference latency, and achieves state-of-the-art accuracy.

Technology Category

Computer Vision: Large Vision ModelsMachine Learning: Hardware-aware MLNatural Language Processing: Sentence-level Semantics, Textual Inference, etc.

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Vertical and domain-specific searchSecurity and Privacy: Privacy-enhancing technologies
📝 Abstract
Parameter-efficient fine-tuning (PEFT) has emerged as a popular solution for adapting pre-trained Vision Transformer (ViT) models to downstream applications. While current PEFT methods have achieved parameter efficiency, they overlook the efficiency of computation and GPU memory during both fine-tuning and inference, falling short of practical requirements. In this paper, we propose extbf{Sparse-Tuning}, a novel PEFT method that accounts for the information redundancy in images and videos to boost the above efficiency. By sparsely preserving the semantic-relevant tokens and merging irrelevant ones, Sparse-Tuning minimizes the quantity of tokens processed at each layer, leading to a quadratic reduction in computational and memory overhead. To align our token sparsification strategy suitably with fine-tuning purposes, we further design Dense Adapters that establish dense connections from shallow layers to deeper layers. These Dense Adapters integrate multi-level local features to enrich the current tokens, improving both token preservation and model adaptation. Empirical results on VTAB-1K, three image datasets, and two video datasets show that our Sparse-Tuning reduces GFLOPs to extbf{62%-70%} of the original ViT-B while achieving state-of-the-art performance. Source code is available at url{https://github.com/liuting20/Sparse-Tuning}.
Problem

Research questions and friction points this paper is trying to address.

Improves inference efficiency of Vision Transformers
Reduces computational and memory overhead in fine-tuning
Maintains performance despite token sparsification information loss
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse-Tuning integrates token sparsification to cut computational costs
Dense Adapters compensate for information loss from token sparsification
The framework reduces GFLOPs to 66% while maintaining top performance
🔎 Similar Papers
No similar papers found.