Memory-Efficient Split Federated Learning for LLM Fine-Tuning on Heterogeneous Mobile Devices

📅 2025-06-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address memory constraints and computational imbalance in fine-tuning large language models (LLMs) on heterogeneous mobile devices, this paper proposes an edge-cooperative lightweight split federated learning framework. The method decomposes the LLM across edge and client tiers, enabling collaborative training under resource limitations. Its key contributions are: (1) deploying LoRA adapters exclusively in bottom-layer modules on clients to minimize local memory overhead; (2) sequentially updating module parameters at the server and introducing a novel hierarchical LoRA allocation mechanism that dynamically matches adapter configurations to device capabilities and memory budgets; and (3) designing a resource-aware dynamic device scheduling algorithm to orchestrate edge-terminal cooperation. Experiments demonstrate that the approach reduces memory consumption by 79% and shortens training time by 6% compared to baselines, while achieving model accuracy comparable to full-parameter fine-tuning.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsComputer Vision: Large Vision Models

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Large language models for searchSystems and Infrastructure for Web, Mobile and WoT: Federated Web and WoT systems, including distributed, federated and edge-based data processing
📝 Abstract
In this paper, we propose an edge-assisted split federated learning framework to facilitate large language model (LLM) fine-tuning on heterogeneous mobile devices while alleviating memory pressures on both mobile devices and the edge server. Specifically, mobile devices perform low-rank adaptation (LoRA) fine-tuning on only a subset of lower layers of the pre-trained LLM, tailored to their individual capacities. On the server, a full LLM is maintained, and the corresponding LoRA modules are selectively fine-tuned in a sequential manner for each device. To further enhance training efficiency, we propose a server-side training scheduling method that optimizes the processing order of devices for accelerating fine-tuning. Extensive experiments demonstrate that compared to the baselines, our scheme can reduce 79% memory footprint and 6% training time while achieving comparable performance.
Problem

Research questions and friction points this paper is trying to address.

Memory-efficient LLM fine-tuning on mobile devices
Edge-assisted split federated learning for heterogeneous devices
Reducing memory footprint and training time
Innovation

Methods, ideas, or system contributions that make the work stand out.

Edge-assisted split federated learning framework
Low-rank adaptation for subset fine-tuning
Server-side training scheduling optimization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Xiaopei Chen
Xiaopei Chen
South China University of Technology
edge intelligencewireless communications
L
Liang Li
Frontier Research Center, Peng Cheng Laboratory, Shenzhen, China
F
Fei Ji
School of Electronic and Information Engineering, South China University of Technology, Guangzhou, China
W
Wen Wu
Frontier Research Center, Peng Cheng Laboratory, Shenzhen, China