RingAda: Pipelining Large Model Fine-Tuning on Edge Devices with Scheduled Layer Unfreezing

📅 2025-02-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Efficient personalized fine-tuning of large language models (LLMs) on edge devices faces three core challenges: severe memory constraints, limited computational resources, and stringent privacy requirements. To address these, we propose RingTune—a distributed fine-tuning framework featuring a ring-topology pipelined parallelism scheme that collaboratively partitions frozen Transformer blocks and lightweight adapters across multiple edge devices. It introduces a hierarchical dynamic unfreezing strategy enabling batch-wise training and top-down progressive parameter activation. Moreover, RingTune pioneers adapter-level gradient early-stopping backward propagation to eliminate redundant computation. Compared to centralized fine-tuning, RingTune reduces GPU memory consumption by 47% and training time by 39%, while retaining 98.6% of downstream task performance. For the first time, RingTune enables low-overhead, high-accuracy, and privacy-preserving continual adaptation of LLMs on edge devices—fully respecting data locality.

Technology Category

Machine Learning: Learning on the Edge & Model CompressionSearch and Optimization: Distributed SearchNatural Language Processing: (Large) Language Models

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
To enable large model (LM) based edge intelligent service provisioning, on-device fine-tuning with locally personalized data allows for continuous and privacy-preserving LM customization. In this paper, we propose RingAda, a collaborative training framework designed for fine-tuning transformer-based LMs on edge devices. Particularly, RingAda performs parameter-efficient adapter fine-tuning across a set of interconnected edge devices, forming a ring topology for per-batch training by sequentially placing frozen transformer blocks and their trainable adapter modules on the devices. RingAda follows a novel pipeline-parallel training mechanism with top-down adapter unfreezing, allowing for early-stopping of backpropagation at the lowest unfrozen adapter layer, thereby accelerating the fine-tuning process. Extensive experimental results demonstrate that RingAda significantly reduces fine-tuning time and memory costs while maintaining competitive model performance compared to its peer designs.
Problem

Research questions and friction points this paper is trying to address.

Fine-tuning large models on edge devices
Reducing memory and time costs
Maintaining model performance efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Pipeline-parallel training mechanism
Parameter-efficient adapter fine-tuning
Top-down adapter unfreezing strategy
💼 Related Jobs
No related jobs found.
L
Liang Li
Frontier Research Center, Peng Cheng Laboratory, Shenzhen, China
Xiaopei Chen
Xiaopei Chen
South China University of Technology
edge intelligencewireless communications
W
Wen Wu
Frontier Research Center, Peng Cheng Laboratory, Shenzhen, China