FedLore: Communication and Memory Efficient Federated Learning via Shared Gradient Low-Rank Projection

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of independent low-rank projections in federated learning, which cause subspace fragmentation, aggregation bias, and substantial memory and communication overhead. We propose a shared gradient low-rank projection framework that achieves exact aggregation in low-rank coordinates through intra-round basis sharing, inter-round dynamic refreshing, and a projected SGD algorithm. This approach theoretically eliminates projection bias and overcomes fixed rank-budget constraints. Experiments demonstrate that the proposed method outperforms LoRA baselines and matches full-parameter fine-tuning on vision-language tasks while significantly reducing communication costs and optimizer state memory.
📝 Abstract
Federated training of foundation models is constrained by client memory and communication costs. LoRA-based methods reduce these costs through low-rank adapters, but their fixed rank budget can limit adaptation. Gradient low-rank optimization offers greater flexibility, yet independently chosen client subspaces create a problem we term \emph{subspace fragmentation}: local projections interact with data heterogeneity to bias aggregated directions, while aggregation can increase update rank and communication cost. Thus, accurate local gradient compression need not preserve global descent. We propose \texttt{FedLore}, which shares a low-rank optimization basis within each round and refreshes it across rounds. The shared basis enables exact aggregation in low-rank coordinates and eliminates the identified projection bias. Subspace refresh allows the accumulated model update to exceed the per-round rank budget. We characterize the aggregation bias and establish an $O(T^{-1/2})$ stationarity bound for the projected-SGD variant under a global-gradient coverage condition and standard smoothness and variance assumptions, with bounded gradient heterogeneity. Experiments on vision and language tasks, including federated pre-training, show that \texttt{FedLore} outperforms the evaluated low-rank adapter baselines and matches or exceeds full-parameter training, while reducing communication and optimizer-state memory.
Problem

Research questions and friction points this paper is trying to address.

Federated Learning
Foundation Models
Low-Rank Optimization
Subspace Fragmentation
Communication Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Federated Learning
Low-Rank Projection
Subspace Fragmentation
Communication Efficiency
Foundation Models
🔎 Similar Papers
No similar papers found.
J
Junkang Liu
Tianjin University