🤖 AI Summary
This study addresses the structural aggregation mismatch and statistical collaboration mismatch caused by data heterogeneity in federated LoRA fine-tuning. To this end, we propose an optimization framework based on decoupled factor aggregation and adaptive clustering. Methodologically, the approach preserves the low-rank structure by sharing the A-factor while personalizing the B-factor, and mitigates negative transfer through adaptation-aware clustering leveraging cosine similarity. Notably, only a single LoRA factor requires communication per round. Experimental results demonstrate that the proposed method achieves the highest average accuracy across multimodal tasks using RoBERTa and ViT architectures. Ultimately, this work enables efficient federated parameter-efficient fine-tuning with minimal communication overhead.
📝 Abstract
Federated LoRA fine-tuning enables parameter-efficient adaptation of pre-trained models without sharing private data, but suffers from two fundamental mismatches under heterogeneous client data: a structural aggregation mismatch caused by independently averaging LoRA factors, and a statistical collaboration mismatch caused by enforcing a single global adapter across divergent clients. To address these issues, we propose CF-LoRA, a clustered federated LoRA fine-tuning framework that combines decoupled factor aggregation with adaptation-aware client clustering. CF-LoRA first learns a globally shared $A$ factor while retaining personalized $B_i$ factors, then identifies clients with similar adaptation patterns based on the cosine similarity of their learned $B_i$ factors, and finally performs intra-cluster $B$-factor aggregation with a frozen $A$ factor. By decoupling LoRA factor aggregation, CF-LoRA preserves the low-rank structure and mitigates the structural aggregation mismatch, while adaptation-aware clustering promotes collaboration among clients with similar adaptation patterns and reduces negative transfer caused by statistical heterogeneity. Experiments on four language tasks and four vision datasets with RoBERTa and ViT show that CF-LoRA achieves the highest average accuracy in both modalities while communicating only one LoRA factor per optimization round.