🤖 AI Summary
To address memory constraints and computational imbalance in fine-tuning large language models (LLMs) on heterogeneous mobile devices, this paper proposes an edge-cooperative lightweight split federated learning framework. The method decomposes the LLM across edge and client tiers, enabling collaborative training under resource limitations. Its key contributions are: (1) deploying LoRA adapters exclusively in bottom-layer modules on clients to minimize local memory overhead; (2) sequentially updating module parameters at the server and introducing a novel hierarchical LoRA allocation mechanism that dynamically matches adapter configurations to device capabilities and memory budgets; and (3) designing a resource-aware dynamic device scheduling algorithm to orchestrate edge-terminal cooperation. Experiments demonstrate that the approach reduces memory consumption by 79% and shortens training time by 6% compared to baselines, while achieving model accuracy comparable to full-parameter fine-tuning.
📝 Abstract
In this paper, we propose an edge-assisted split federated learning framework to facilitate large language model (LLM) fine-tuning on heterogeneous mobile devices while alleviating memory pressures on both mobile devices and the edge server. Specifically, mobile devices perform low-rank adaptation (LoRA) fine-tuning on only a subset of lower layers of the pre-trained LLM, tailored to their individual capacities. On the server, a full LLM is maintained, and the corresponding LoRA modules are selectively fine-tuned in a sequential manner for each device. To further enhance training efficiency, we propose a server-side training scheduling method that optimizes the processing order of devices for accelerating fine-tuning. Extensive experiments demonstrate that compared to the baselines, our scheme can reduce 79% memory footprint and 6% training time while achieving comparable performance.