federated optimizer selection

Designs, builds, and evaluates optimization algorithms and strategies for federated learning training, including the implementation and comparison of methods (e.g., FedAvg, FedProx, adaptive federated optimizers) and their hyperparameter choices. Analyzes optimizer behavior to select approaches that ensure convergence, robustness to client heterogeneity, and acceptable communication and computation trade-offs in decentralized training.

federatedoptimizerselection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.2
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Parameter Tracking in Federated Learning with Adaptive Optimization

Feb 04, 2025
EC
Evan Chen
🏛️ Purdue University | IBM

Severe data heterogeneity across clients in federated learning severely degrades model convergence, and existing gradient tracking (GT) methods are limited to SGD, lacking compatibility with mainstream adaptive optimizers such as Adam. Method: We propose a novel *parameter tracking* (PT) paradigm that generalizes GT from the gradient space to the parameter space—enabling, for the first time, tight integration with Adam. Based on PT, we design two new federated adaptive algorithms: FAdamGT and FAdamET. Contribution/Results: Theoretically, we provide the first rigorous convergence guarantee for adaptive federated optimization under non-convex objectives. Technically, we achieve this via distributed first-order information correction and a principled federated adaptation of Adam, balancing communication efficiency and convergence stability. Extensive experiments demonstrate that our methods significantly reduce both communication and computational overhead across diverse heterogeneity settings, consistently outperforming state-of-the-art federated SGD and adaptive baselines.

Address data heterogeneity in Federated LearningEnhance convergence with Parameter Tracking in FLExtend Gradient Tracking to adaptive optimizers

Distributed optimization: designed for federated learning

Aug 11, 2025
WG
Wenyou Guo
🏛️ Jinan University | Guangdong International Cooperation Base of Science and Technology for GBA Smart Logistics | School of Intelligent Systems Science and Engineering | Institute of Physical Internet | Jiangxi University of Science and Technology | The Hong Kong Polytechnic University

To address the challenges of high statistical heterogeneity, diverse communication topologies, and stringent privacy constraints in cross-organizational federated learning (FL), this paper proposes a generic distributed optimization algorithm grounded in the augmented Lagrangian framework. Methodologically, it integrates proximal relaxation with quadratic approximation techniques, enabling unified convergence analysis for variants including proximal gradient descent and stochastic gradient descent—supporting both centralized and decentralized topologies, asynchronous communication, and non-IID data. Theoretically, it establishes the first general convergence analysis framework compatible with multiple FL architectures and termination criteria. Empirically, the algorithm achieves significantly faster convergence and improved communication efficiency in large-scale, highly heterogeneous settings, while demonstrating robustness and practical applicability.

Develop distributed optimization for federated learning privacy constraintsEnhance efficiency with termination criteria and convergence guaranteesPropose algorithms for diverse FL communication topologies

MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters

Feb 04, 2024
AS
Arsalan Sharifnassab
🏛️ University of Alberta | Leiden University

To address the inefficiency and poor generalizability of manual hyperparameter tuning—particularly for learning rates—this paper proposes a dynamic online meta-optimization framework that formulates learning rate adaptation as a discounted cumulative regret minimization problem over time. The method employs a gradient-based meta-update mechanism, enabling plug-and-play integration with any first-order optimizer (e.g., SGD, Adam) to achieve decoupled, real-time, adaptive step-size optimization. Key contributions include: (i) the first formalization of meta-optimization as discounted regret minimization; and (ii) a low-complexity variant that preserves theoretical rigor while ensuring computational efficiency and strong generalization. Experiments across diverse tasks demonstrate faster convergence, enhanced robustness to initialization and task heterogeneity, competitive performance against hand-tuned optimal schedulers, and significantly lower computational overhead compared to conventional hyperparameter search methods.

Dynamically adjusting step sizes during model optimizationOptimizing meta-parameters for efficient machine learning trainingReducing regret by considering long-term impact of learning rates

This paper addresses the inefficiency and lack of scalability of manual hyperparameter tuning in large-scale machine learning. It systematically surveys hyperparameter optimization (HPO), unifying and classifying five mainstream paradigms: random/low-discrepancy search, bandit-based methods, Bayesian optimization, population-based (evolutionary) algorithms, and gradient-based differentiable optimization. The survey further extends to emerging settings—including online HPO, constrained HPO, and multi-objective HPO. Crucially, the work establishes novel theoretical connections between HPO and meta-learning as well as neural architecture search, yielding a comprehensive knowledge framework that articulates methodological principles, applicability boundaries, and inherent limitations. By clarifying the technical evolution and identifying key open challenges, this study provides a theoretically grounded yet practically actionable foundation for automated machine learning.

Addressing challenges in online, constrained, and multi-objective hyperparameter tuningAutomating hyperparameter search to improve machine learning efficiencyComparing state-of-the-art hyperparameter optimization techniques and methods

Exploring Parameter-Efficient Fine-Tuning to Enable Foundation Models in Federated Learning

Oct 04, 2022
GS
Guangyu Sun
🏛️ University of Central Florida

How can large pre-trained models be efficiently deployed in federated learning while balancing communication efficiency and model performance? This paper introduces FedPEFT, the first systematic integration of parameter-efficient fine-tuning (PEFT) into federated learning. In FedPEFT, clients update only a small set of trainable modules—e.g., LoRA or Adapters—while the server performs lightweight aggregation of these sparse updates. The framework natively accommodates practical constraints including Non-IID data distributions, client dropouts, and differential privacy requirements. Extensive experiments across multiple federated benchmarks demonstrate that FedPEFT reduces total communication overhead by up to 95% compared to standard baselines, while matching or surpassing the accuracy of FedAvg. These results significantly enhance the practical feasibility of deploying large language models in resource-constrained edge environments.

Communication EfficiencyFederated LearningPre-trained Models

Latest Papers

What's happening recently
View more

This paper identifies the fundamental mechanism behind performance degradation in federated optimization under data heterogeneity: discrepancies among clients’ local optima elevate the lower bound of the global objective function, rendering perfect global fit infeasible and causing the global model to converge to an oscillatory region rather than a fixed point. Method: Grounded in distributed optimization theory, we establish the first rigorous analytical link between local optimal divergence and global convergence behavior. Our approach integrates theoretical derivation with empirical validation across diverse tasks and model architectures, and we open-source a unified framework, FedTorch. Contribution/Results: We provide a verifiable theoretical explanation for federated learning’s performance degradation. We prove—both theoretically and empirically—that global models cannot perfectly fit all client data under heterogeneity, and that convergence oscillation is an inherent, provable phenomenon. This work offers a novel theoretical perspective and a testable foundation for federated optimization.

Demonstrates global model oscillation limits convergence to optimumExplains performance degradation in federated learning under data heterogeneityShows client-side optima prevent perfect fitting of all data

This work addresses the client drift problem in federated learning caused by non-independent and identically distributed (non-IID) data by proposing FedZMG, an optimization algorithm that incurs no additional parameters or communication overhead. FedZMG mitigates gradient bias induced by data heterogeneity by projecting local gradients onto a zero-mean hyperplane, thereby structurally regularizing the optimization space. Theoretical analysis demonstrates that FedZMG reduces gradient variance and improves convergence guarantees. Extensive experiments on highly non-IID benchmarks—including EMNIST, CIFAR100, and Shakespeare—show that FedZMG consistently outperforms FedAvg and FedAdam, achieving faster convergence and higher final validation accuracy without increasing computational or communication costs.

client driftconvergence degradationFederated Learning

To address the client drift and generalization imbalance in federated learning under non-independent and identically distributed (Non-IID) data, this paper identifies a critical limitation of existing personalization methods: their excessive focus on local accuracy while neglecting out-of-distribution (OOD) generalization—a fundamental pillar of FedAvg’s robustness. We propose a unified evaluation paradigm that jointly optimizes local accuracy and OOD generalization, and design FLIU, an adaptive personalization update mechanism. Within the FedAvg framework, FLIU introduces learnable, client-specific scaling factors to dynamically balance global consistency and local adaptability. Extensive experiments across MNIST and CIFAR-10 under IID, pathological Non-IID, and Dirichlet Non-IID settings demonstrate that FLIU achieves high local accuracy while significantly improving OOD generalization—outperforming state-of-the-art personalized federated learning methods.

Addressing client drift in federated learning with heterogeneous data distributionsEvaluating local performance versus out-of-distribution generalization in personalized FLProposing individualized updates to improve both local and global model performance

FedPM: Federated Learning Using Second-order Optimization with Preconditioned Mixing of Local Parameters

Nov 12, 2025
HI
Hiro Ishii
🏛️ Institute of Science Tokyo | NTT Communication Science Laboratories

Existing federated learning methods (e.g., LocalNewton, LTDA, FedSophia) suffer from slow convergence under data heterogeneity due to drift of local preconditioners. This work proposes FedPM, the first framework to introduce a preconditioner parameter mixing mechanism at the server side. FedPM decomposes the global second-order update into two components: gradient preconditioning and local update correction—thereby fundamentally mitigating preconditioner drift. We establish theoretical guarantees showing that FedPM achieves superlinear convergence under strong convexity. Extensive experiments on multiple heterogeneous benchmarks demonstrate significant improvements in test accuracy, validating both the efficacy and stability of second-order optimization in federated learning. The core contributions are (i) a novel server-side preconditioner mixing paradigm and (ii) rigorous convergence analysis ensuring superlinear rates under standard assumptions.

Enhancing test accuracy through preconditioned parameter mixing on serverImproving convergence in heterogeneous data settings with second-order optimizationMitigating local preconditioner drift in federated learning systems

This work addresses the challenges of data privacy and communication efficiency in customizing large models within federated learning. The authors systematically evaluate the suitability of various parameter-efficient fine-tuning methods and, for the first time, introduce prefix-tuning into the federated learning framework, proposing Federated Prefix-Tuning. This approach achieves model performance comparable to centralized training while significantly improving communication efficiency and robustness. Experimental results across multiple tasks demonstrate that the proposed method either outperforms or matches existing federated customization strategies—including full fine-tuning, other parameter-efficient fine-tuning techniques, and knowledge distillation—thereby validating its effectiveness and practicality.

distributed optimizationfederated learninglarge model customization

Hot Scholars

FS

Fanhua Shang

Professor at Tianjin University
Machine LearningData MiningComputer Vision
HL

Hongying Liu

Tianjin University
Machine learningImage processing
JK

Joonhyuk Kang

Professor of Electrical Engineering, KAIST
Signal Processing and Machine Learning for Wireless Communication Systems
YL

Youngjoon Lee

Ph.D. student, KAIST
machine learningdistributed learningprivacy-preserving
PV

Praneeth Vepakomma

Massachusetts Institute of Technology, MBZUAI
Collaborative MLResponsible AIPrivacy