FedFit: Federated Fine-Tuning of LLMs via Vector-Bank Parameterization and Quantization

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the substantial communication overhead and LoRA aggregation conflicts inherent in federated fine-tuning of large language models by proposing a parameterized reconstruction method based on a decoupled shared vector library. Through an alternating optimization strategy and a residual spectral aggregation mechanism, the approach effectively resolves Sum-of-Products (SoP) versus Product-of-Sums (PoS) conflicts. Furthermore, it integrates block quantization with error feedback to compress transmitted data, thereby enabling efficient and privacy-preserving fine-tuning. Experiments conducted on Qwen2.5 demonstrate that the proposed method achieves over a hundredfold communication compression ratio while yielding perplexity performance comparable to standard centralized training. Additionally, the framework is supported by rigorous convergence guarantees.
📝 Abstract
Federated Learning (FL) enables privacy-preserving fine-tuning of Large Language Models (LLMs), yet the massive communication overhead remains a critical bottleneck. Furthermore, applying Low-Rank Adaptation (LoRA) in FL faces a fundamental "aggregation dilemma" between the accurate Sum-of-Products (SoP) and the communication-efficient Product-of-Sums (PoS) implementations. To tackle these challenges, we propose FedFit. First, to significantly reduce communication overhead, we introduce a disjoint shared vector-bank parameterization that reconstructs high-dimensional adapter matrices from two compact and disjoint global vector banks. Second, to address the aggregation dilemma, we devise an alternating optimization schedule. By cycling between decoupled single-bank updates (which allow for accurate aggregation) and joint updates corrected by a Residual Spectral Aggregation mechanism, we resolve the conflict between SoP and PoS. Additionally, we integrate blockwise quantization with client-side error feedback to further compress the transmitted vectors. Furthermore, we establish theoretical convergence guarantees for the proposed algorithm. Extensive experiments on Qwen2.5 models demonstrate that FedFit achieves perplexity performance comparable to standard federated LoRA methods, while providing compression ratios up to 100x higher.
Problem

Research questions and friction points this paper is trying to address.

Federated Learning
Large Language Models
Fine-Tuning
Communication Overhead
Low-Rank Adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Federated Learning
Low-Rank Adaptation
Vector-Bank Parameterization
Alternating Optimization
Quantization
H
Hang Zou
6G Research Center, Khalifa University, 127788 Abu Dhabi, UAE
C
Chao Zhang
Department of Computer Science, Central South University, 410083, Changsha, China
Y
Yuzhi Yang
6G Research Center, Khalifa University, 127788 Abu Dhabi, UAE
Y
Yu Tian
6G Research Center, Khalifa University, 127788 Abu Dhabi, UAE
Samson Lasaulce
Samson Lasaulce
CNRS Director of Research
Game theoryOptimizationLearningNetworks
M
Mérouane Debbah
6G Research Center, Khalifa University, 127788 Abu Dhabi, UAE