🤖 AI Summary
This study addresses the aggregation error introduced by separately averaging low-rank factors in federated LoRA, a challenge that existing methods fail to reconcile with simultaneous factor updates and zero-error aggregation. To overcome this limitation, this work proposes CFLoRA, which leverages complementary partitioned latent channels to eliminate bilinear terms. This approach achieves, for the first time, error-free aggregation with synchronous factor updates while supporting heterogeneous rank budgets. Theoretically, we establish a convergence rate of O(1/√T). Empirically, evaluations on GLUE and commonsense reasoning benchmarks demonstrate that CFLoRA significantly outperforms state-of-the-art baselines in both model performance and training efficiency.
📝 Abstract
Federated low-rank adaptation (LoRA) enables collaborative fine-tuning of large language models without centralizing private client data. Its factorized update, however, creates a structural mismatch in federated averaging: averaging the two LoRA factors separately does not equal averaging their products. Existing exact methods resolve this issue mainly by freezing an entire factor or alternating factors across rounds, but none can update factors simultaneously without aggregation errors or expanding communication ranks. To address this fundamental problem, we present \texttt{CFLoRA}, a federated LoRA scheme that partitions latent LoRA channels into two complementary sets in every communication round. By ensuring that columns and rows are complementary across factors, we eliminate bilinear terms in matrix multiplications, making federated aggregation exact. Crucially, our framework also supports clients with heterogeneous rank budgets. Convergence analysis validates \texttt{CFLoRA} achieves $\mathcal{O}(1/\sqrt{T})$ convergence rate of the \textit{original} LoRA objective in homogeneous-rank cases. Extensive experiments with RoBERTa on the GLUE benchmark and with LLaMA-3.2-3B-Instruct on commonsense reasoning tasks demonstrate that \texttt{CFLoRA} achieves superior performance and training efficiency compared to state-of-the-art federated LoRA baselines.