Score
Designs, builds, and evaluates optimization algorithms and strategies for federated learning training, including the implementation and comparison of methods (e.g., FedAvg, FedProx, adaptive federated optimizers) and their hyperparameter choices. Analyzes optimizer behavior to select approaches that ensure convergence, robustness to client heterogeneity, and acceptable communication and computation trade-offs in decentralized training.
Severe data heterogeneity across clients in federated learning severely degrades model convergence, and existing gradient tracking (GT) methods are limited to SGD, lacking compatibility with mainstream adaptive optimizers such as Adam. Method: We propose a novel *parameter tracking* (PT) paradigm that generalizes GT from the gradient space to the parameter space—enabling, for the first time, tight integration with Adam. Based on PT, we design two new federated adaptive algorithms: FAdamGT and FAdamET. Contribution/Results: Theoretically, we provide the first rigorous convergence guarantee for adaptive federated optimization under non-convex objectives. Technically, we achieve this via distributed first-order information correction and a principled federated adaptation of Adam, balancing communication efficiency and convergence stability. Extensive experiments demonstrate that our methods significantly reduce both communication and computational overhead across diverse heterogeneity settings, consistently outperforming state-of-the-art federated SGD and adaptive baselines.
To address the challenges of high statistical heterogeneity, diverse communication topologies, and stringent privacy constraints in cross-organizational federated learning (FL), this paper proposes a generic distributed optimization algorithm grounded in the augmented Lagrangian framework. Methodologically, it integrates proximal relaxation with quadratic approximation techniques, enabling unified convergence analysis for variants including proximal gradient descent and stochastic gradient descent—supporting both centralized and decentralized topologies, asynchronous communication, and non-IID data. Theoretically, it establishes the first general convergence analysis framework compatible with multiple FL architectures and termination criteria. Empirically, the algorithm achieves significantly faster convergence and improved communication efficiency in large-scale, highly heterogeneous settings, while demonstrating robustness and practical applicability.
To address the inefficiency and poor generalizability of manual hyperparameter tuning—particularly for learning rates—this paper proposes a dynamic online meta-optimization framework that formulates learning rate adaptation as a discounted cumulative regret minimization problem over time. The method employs a gradient-based meta-update mechanism, enabling plug-and-play integration with any first-order optimizer (e.g., SGD, Adam) to achieve decoupled, real-time, adaptive step-size optimization. Key contributions include: (i) the first formalization of meta-optimization as discounted regret minimization; and (ii) a low-complexity variant that preserves theoretical rigor while ensuring computational efficiency and strong generalization. Experiments across diverse tasks demonstrate faster convergence, enhanced robustness to initialization and task heterogeneity, competitive performance against hand-tuned optimal schedulers, and significantly lower computational overhead compared to conventional hyperparameter search methods.
This paper addresses the inefficiency and lack of scalability of manual hyperparameter tuning in large-scale machine learning. It systematically surveys hyperparameter optimization (HPO), unifying and classifying five mainstream paradigms: random/low-discrepancy search, bandit-based methods, Bayesian optimization, population-based (evolutionary) algorithms, and gradient-based differentiable optimization. The survey further extends to emerging settings—including online HPO, constrained HPO, and multi-objective HPO. Crucially, the work establishes novel theoretical connections between HPO and meta-learning as well as neural architecture search, yielding a comprehensive knowledge framework that articulates methodological principles, applicability boundaries, and inherent limitations. By clarifying the technical evolution and identifying key open challenges, this study provides a theoretically grounded yet practically actionable foundation for automated machine learning.
How can large pre-trained models be efficiently deployed in federated learning while balancing communication efficiency and model performance? This paper introduces FedPEFT, the first systematic integration of parameter-efficient fine-tuning (PEFT) into federated learning. In FedPEFT, clients update only a small set of trainable modules—e.g., LoRA or Adapters—while the server performs lightweight aggregation of these sparse updates. The framework natively accommodates practical constraints including Non-IID data distributions, client dropouts, and differential privacy requirements. Extensive experiments across multiple federated benchmarks demonstrate that FedPEFT reduces total communication overhead by up to 95% compared to standard baselines, while matching or surpassing the accuracy of FedAvg. These results significantly enhance the practical feasibility of deploying large language models in resource-constrained edge environments.
This paper identifies the fundamental mechanism behind performance degradation in federated optimization under data heterogeneity: discrepancies among clients’ local optima elevate the lower bound of the global objective function, rendering perfect global fit infeasible and causing the global model to converge to an oscillatory region rather than a fixed point. Method: Grounded in distributed optimization theory, we establish the first rigorous analytical link between local optimal divergence and global convergence behavior. Our approach integrates theoretical derivation with empirical validation across diverse tasks and model architectures, and we open-source a unified framework, FedTorch. Contribution/Results: We provide a verifiable theoretical explanation for federated learning’s performance degradation. We prove—both theoretically and empirically—that global models cannot perfectly fit all client data under heterogeneity, and that convergence oscillation is an inherent, provable phenomenon. This work offers a novel theoretical perspective and a testable foundation for federated optimization.
This work addresses the client drift problem in federated learning caused by non-independent and identically distributed (non-IID) data by proposing FedZMG, an optimization algorithm that incurs no additional parameters or communication overhead. FedZMG mitigates gradient bias induced by data heterogeneity by projecting local gradients onto a zero-mean hyperplane, thereby structurally regularizing the optimization space. Theoretical analysis demonstrates that FedZMG reduces gradient variance and improves convergence guarantees. Extensive experiments on highly non-IID benchmarks—including EMNIST, CIFAR100, and Shakespeare—show that FedZMG consistently outperforms FedAvg and FedAdam, achieving faster convergence and higher final validation accuracy without increasing computational or communication costs.
To address the client drift and generalization imbalance in federated learning under non-independent and identically distributed (Non-IID) data, this paper identifies a critical limitation of existing personalization methods: their excessive focus on local accuracy while neglecting out-of-distribution (OOD) generalization—a fundamental pillar of FedAvg’s robustness. We propose a unified evaluation paradigm that jointly optimizes local accuracy and OOD generalization, and design FLIU, an adaptive personalization update mechanism. Within the FedAvg framework, FLIU introduces learnable, client-specific scaling factors to dynamically balance global consistency and local adaptability. Extensive experiments across MNIST and CIFAR-10 under IID, pathological Non-IID, and Dirichlet Non-IID settings demonstrate that FLIU achieves high local accuracy while significantly improving OOD generalization—outperforming state-of-the-art personalized federated learning methods.
Existing federated learning methods (e.g., LocalNewton, LTDA, FedSophia) suffer from slow convergence under data heterogeneity due to drift of local preconditioners. This work proposes FedPM, the first framework to introduce a preconditioner parameter mixing mechanism at the server side. FedPM decomposes the global second-order update into two components: gradient preconditioning and local update correction—thereby fundamentally mitigating preconditioner drift. We establish theoretical guarantees showing that FedPM achieves superlinear convergence under strong convexity. Extensive experiments on multiple heterogeneous benchmarks demonstrate significant improvements in test accuracy, validating both the efficacy and stability of second-order optimization in federated learning. The core contributions are (i) a novel server-side preconditioner mixing paradigm and (ii) rigorous convergence analysis ensuring superlinear rates under standard assumptions.
This work addresses the challenges of data privacy and communication efficiency in customizing large models within federated learning. The authors systematically evaluate the suitability of various parameter-efficient fine-tuning methods and, for the first time, introduce prefix-tuning into the federated learning framework, proposing Federated Prefix-Tuning. This approach achieves model performance comparable to centralized training while significantly improving communication efficiency and robustness. Experimental results across multiple tasks demonstrate that the proposed method either outperforms or matches existing federated customization strategies—including full fine-tuning, other parameter-efficient fine-tuning techniques, and knowledge distillation—thereby validating its effectiveness and practicality.