🤖 AI Summary
This work addresses the client drift problem in federated learning caused by non-independent and identically distributed (non-IID) data by proposing FedZMG, an optimization algorithm that incurs no additional parameters or communication overhead. FedZMG mitigates gradient bias induced by data heterogeneity by projecting local gradients onto a zero-mean hyperplane, thereby structurally regularizing the optimization space. Theoretical analysis demonstrates that FedZMG reduces gradient variance and improves convergence guarantees. Extensive experiments on highly non-IID benchmarks—including EMNIST, CIFAR100, and Shakespeare—show that FedZMG consistently outperforms FedAvg and FedAdam, achieving faster convergence and higher final validation accuracy without increasing computational or communication costs.
📝 Abstract
Federated Learning (FL) enables distributed model training on edge devices while preserving data privacy. However, clients tend to have non-Independent and Identically Distributed (non-IID) data, which often leads to client-drift, and therefore diminishing convergence speed and model performance. While adaptive optimizers have been proposed to mitigate these effects, they frequently introduce computational complexity or communication overhead unsuitable for resource-constrained IoT environments. This paper introduces Federated Zero Mean Gradients (FedZMG), a novel, parameter-free, client-side optimization algorithm designed to tackle client-drift by structurally regularizing the optimization space. Advancing the idea of Gradient Centralization, FedZMG projects local gradients onto a zero-mean hyperplane, effectively neutralizing the"intensity"or"bias"shifts inherent in heterogeneous data distributions without requiring additional communication or hyperparameter tuning. A theoretical analysis is provided, proving that FedZMG reduces the effective gradient variance and guarantees tighter convergence bounds compared to standard FedAvg. Extensive empirical evaluations on EMNIST, CIFAR100, and Shakespeare datasets demonstrate that FedZMG achieves better convergence speed and final validation accuracy compared to the baseline FedAvg and the adaptive optimizer FedAdam, particularly in highly non-IID settings.