Minimal Interaction Separated Tuning: A New Paradigm for Visual Adaptation

📅 2024-06-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address three key bottlenecks in fine-tuning large vision models on resource-constrained edge devices—poor adaptability, parameter-heavy adapters, and excessive communication overhead—this paper proposes a cloud-edge decoupled fine-tuning paradigm: the pre-trained visual backbone remains fixed on the cloud for feature extraction, while only a lightweight attention adapter is optimized on-device. We introduce a “minimal interaction” mechanism that transmits only the sum of aggregated, compressed intermediate features—drastically reducing communication volume. Our approach achieves efficient co-optimization across model parameters, FLOPs, memory footprint, and communication bandwidth. Evaluated on multiple vision adaptation benchmarks, it attains state-of-the-art accuracy while reducing on-device FLOPs and GPU memory usage by over 60% and cutting communication volume by more than 90%, thereby delivering both high performance and strong deployment feasibility.

Technology Category

Machine Learning: Learning on the Edge & Model CompressionComputer Vision: Large Vision ModelsSearch and Optimization: Learning to Search

Application Category

User Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendationSystems and Infrastructure for Web, Mobile and WoT: Cloud, edge and content delivery systems for the WebWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
The rapid scaling of large vision pretrained models makes fine-tuning tasks more and more difficult on devices with low computational resources. We explore a new visual adaptation paradigm called separated tuning, which treats large pretrained models as standalone feature extractors that run on powerful cloud servers. The fine-tuning carries out on devices which possess only low computational resources (slow CPU, no GPU, small memory, etc.) Existing methods that are potentially suitable for our separated tuning paradigm are discussed. But, three major drawbacks hinder their application in separated tuning: low adaptation capability, large adapter network, and in particular, high information transfer overhead. To address these issues, we propose Minimal Interaction Separated Tuning, or MIST, which reveals that the sum of intermediate features from pretrained models not only has minimal information transfer but also has high adaptation capability. With a lightweight attention-based adaptor network, MIST achieves information transfer efficiency, parameter efficiency, computational and memory efficiency, and at the same time demonstrates competitive results on various visual adaptation benchmarks.
Problem

Research questions and friction points this paper is trying to address.

Enables fine-tuning on low-resource devices using cloud-based feature extractors
Reduces information transfer overhead in separated tuning paradigms
Improves adaptation capability with lightweight attention-based adapters
Innovation

Methods, ideas, or system contributions that make the work stand out.

Separated tuning for low-resource devices
Lightweight attention-based adaptor network
Sum of intermediate features minimizes transfer
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Nanjing University
N
Ningyuan Tang
National Key Laboratory for Novel Software Technology, School of Artificial Intelligence, Nanjing University, China
M
Minghao Fu
National Key Laboratory for Novel Software Technology, School of Artificial Intelligence, Nanjing University, China
J
Jianxin Wu
National Key Laboratory for Novel Software Technology, School of Artificial Intelligence, Nanjing University, China