asynchronous block coordinate descent

Designs, implements, and analyzes optimization algorithms and distributed update rules that perform block-coordinate descent asynchronously across multiple processors or agents, including robust variants that tolerate heterogeneous communication and computation delays and permit orthogonal (non-overlapping) block updates. This work covers building asynchronous/distributed implementations that can operate over sliding estimation windows, specifying update schedules and conflict-avoidance mechanisms, and proving convergence properties (e.g., exponential rates) under realistic delay and partial-update models.

asynchronousblockcoordinatedescent

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Asynchronous Decentralized SGD under Non-Convexity: A Block-Coordinate Descent Framework

May 15, 2025
YZ
Yijie Zhou
🏛️ The Chinese University of Hong Kong

Decentralized optimization over heterogeneous devices faces challenges from computational heterogeneity and unpredictable communication delays. Method: We propose Asynchronous Decentralized Stochastic Gradient Descent (ADSGD), a center-free algorithm that eliminates the need for centralized coordination. Contribution/Results: We establish, for the first time, its convergence guarantee for non-convex objectives without assuming bounded data heterogeneity. To enable step-size design independent of computation-communication delays, we develop an Asynchronous Stochastic Block Coordinate Descent (ASBCD) analytical framework. Experiments demonstrate that ADSGD significantly reduces wall-clock time, lowers memory and communication overhead, and exhibits strong robustness to both computational and communication delays—making it well-suited for real-world distributed learning settings.

Addresses challenges in decentralized optimization due to heterogeneous computation speedsEnsures convergence without bounded data heterogeneity assumptionProposes Asynchronous Decentralized SGD resilient to communication delays

This work re-examines the necessity of asynchronous stochastic gradient descent (SGD) in distributed optimization by rigorously analyzing synchronous SGD and its robust variant, m-synchronous SGD, under a realistic heterogeneous setting that combines stochastic computation delays with adversarial partial participation. Through theoretical analysis, the authors demonstrate that the time complexity of these synchronous methods deviates from the optimal only by a logarithmic factor. By introducing a more practical heterogeneity model that unifies both random and adversarial elements, the study systematically evaluates the convergence efficiency of synchronous approaches and reveals their near-optimality across a broad range of heterogeneous scenarios. These findings challenge the prevailing assumption that asynchronous methods are indispensable for efficient distributed optimization in heterogeneous environments.

Asynchronous SGDDistributed OptimizationHeterogeneous Computation

Distributed stochastic optimization faces arbitrary computational dynamics—including hardware disconnections, time-varying compute capacity, and fluctuating processing speeds—rendering existing models inadequate for real-world deployment. Method: We propose the first general asynchronous computation model encompassing all realistic scenarios; based on it, we derive tight time-complexity lower bounds (up to constant factors) for mainstream synchronous and asynchronous methods—including Minibatch SGD, Async SGD, and Picky SGD—and prove that Rennala/Malenia SGD achieves optimal convergence. Our analysis integrates general computational dynamical modeling, information-theoretic lower-bound derivation, and stochastic optimization convergence theory. Contribution: We establish fundamental theoretical limits and design principles for system-aware optimization, providing foundational support for fault-tolerant, robust distributed learning. The results unify treatment of heterogeneous, unreliable, and dynamic execution environments while delivering precise complexity characterizations grounded in both system behavior and statistical learning theory.

Addresses hardware and network delays in parallel computationEstablishes optimal time complexities in distributed stochastic optimizationProves tight lower bounds for synchronous and asynchronous methods

To address low training efficiency, high data/time costs, and severe policy staleness in team reinforcement learning under heterogeneous computing environments, this paper proposes the Asynchronous Federated Policy Gradient (AFedPG) framework. AFedPG enables collaborative training of a global policy across (N) heterogeneous agents and introduces the first lookahead mechanism tailored for federated reinforcement learning (FedRL), which adaptively compensates for policy staleness induced by asynchronous updates. It establishes the first global convergence theory for asynchronous FedRL, yielding an (O(varepsilon^{-2.5}/N)) sample complexity bound and an (Oig((sum 1/t_i)^{-1}ig)) time complexity bound—both demonstrating linear speedup in (N). Experiments on four MuJoCo benchmark tasks show that AFedPG significantly outperforms existing baselines, achieving substantial improvements in wall-clock time efficiency under computational heterogeneity. Theoretical guarantees align closely with empirical results.

Heterogeneous AgentsScalability OptimizationTeam Reinforcement Learning

On the Geometric Convergence of Byzantine-Resilient Distributed Optimization Algorithms

May 18, 2023
KK
Kananart Kuwaranancharoen
🏛️ Purdue University

This paper studies Byzantine-resilient distributed optimization: enabling honest nodes to collaboratively converge to the mean of global optima despite adversarial node attacks. We propose a general algorithmic framework integrating gradient projection, redundancy-based verification, and coordinate-wise robust aggregation (e.g., coordinate-wise median or trimmed mean). We establish, for the first time, a unified theoretical analysis guaranteeing geometric convergence under mild conditions. Specifically, we quantitatively characterize the radius of the ultimate convergence ball in terms of strong convexity, Lipschitz continuity, step size, and network diameter. Theoretically, under weak connectivity and bounded Byzantine fraction, all honest nodes converge geometrically to a neighborhood of the global optimum while achieving approximate consensus. Crucially, the asymptotic error bound is explicitly expressed as a function of network size and Byzantine node proportion, admitting closed-form characterization.

Cost minimizationNetwork dynamicsOptimization under interference

Latest Papers

What's happening recently
View more

This work addresses the challenge of achieving high accuracy and efficiency in distributed trajectory estimation for multi-robot systems under asynchronous communication and computation. The authors propose an asynchronous block coordinate descent algorithm based on sliding-window maximum a posteriori estimation. By incorporating a controllable approximation model, the method reduces communication overhead by up to 96.9% while introducing only negligible estimation error. Theoretical analysis establishes, for the first time, that the algorithm converges exponentially to the optimal solution under asynchronous iterations. Experimental results demonstrate that the proposed approach reduces estimation error by up to 64% compared to state-of-the-art methods and exhibits strong robustness across real-world robotic systems subjected to communication delays spanning three orders of magnitude.

asynchronouscommunication delaysdistributed trajectory estimation

This work addresses the challenge of stochastic gradient optimization in federated learning under delayed and biased gradient information. The authors propose a distributed stochastic optimization framework that enables multiple local agents to collaboratively minimize a global objective despite the presence of delays and biased gradient estimates. Their approach employs a delayed stochastic gradient descent (SGD) algorithm with a predetermined decaying stepsize schedule, deliberately avoiding delay-adaptive mechanisms. Theoretical analysis demonstrates that this strategy achieves the optimal SGD convergence rates in both non-convex and strongly convex settings, thereby establishing the efficacy and superiority of fixed decaying stepsizes in delayed environments.

Diminishing Step SizeDistributed Stochastic OptimizationFederated Learning

In heterogeneous data and system environments, asynchronous stochastic gradient descent (ASGD) biases optimization toward a frequency-weighted average of local objectives due to faster workers updating more frequently, thereby deviating from the true global optimum. This work proposes scaling each worker’s learning rate inversely proportional to its computation time, ensuring that all workers contribute equal total learning rate per unit time—without altering the standard ASGD mechanism or requiring synchronization, buffering, or additional memory. Theoretically, this approach achieves, for the first time in a fixed-computation model under non-convex settings, unbiased convergence to the correct global objective, with the leading term matching the known lower bound on time complexity; asynchronous delays and data heterogeneity affect only lower-order terms. Experiments confirm convergence to the correct solution and demonstrate performance competitive with or superior to state-of-the-art methods.

asynchronous SGDdata heterogeneitydistributed optimization

This work addresses the challenge of gradient staleness in asynchronous stochastic gradient descent caused by data-dependent delays, which existing methods mitigate at the cost of introducing systematic bias through discarding or downweighting delayed gradients. The paper proposes a momentum-based asynchronous optimization framework that preserves information from delayed gradients while effectively alleviating staleness. Under standard assumptions and accounting for data-dependent delays, the method establishes, for the first time, optimal convergence rates for both convex and smooth non-convex optimization problems. Furthermore, it introduces a robust adaptive learning rate scheduling strategy that substantially simplifies hyperparameter tuning. Together, the theoretical analysis and algorithmic design offer a novel analytical perspective and practical tools for asynchronous optimization.

asynchronous SGDconvergence ratesdata-dependent delays

This work addresses the convergence challenges of asynchronous adaptive first-order methods in non-convex stochastic optimization by proposing a class of parallel asynchronous adaptive algorithms that support momentum and inexact normalization, encompassing asynchronous variants of several mainstream optimizers. Under a fully stochastic setting, the paper establishes—for the first time—an $O(1/\sqrt{t})$ convergence rate (up to logarithmic factors) for such methods on non-convex objectives. The theoretical analysis rigorously integrates techniques from asynchronous parallel computation, adaptive learning rates, and stochastic optimization to prove convergence guarantees. Empirical evaluations further demonstrate the algorithm’s efficiency and practicality in heterogeneous large-scale machine learning systems.

adaptive first-order methodsasynchronous optimizationnon-convex optimization

Hot Scholars

NC

Nilanjana Chakraborty

Postdoctoral Researcher, University of Pennsylvania
Bayesian Machine LearningHigh-Dimensional Time SeriesFunctional Data AnalysisCausal Inference
AM

Ankur Mali

Assistant Professor, University of South Florida
Formal languageMemory NetworksPredictive CodingNatural Language Processing
ZW

Zixin Wang

Research Assistant Professor, Department of ECE, The Hong Kong University of Science and Technology
Edge intelligencefederated learningedge large AI modelFoundation Models
SS

Shenghui Song

The Hong Kong University of Science and Technology
Information TheoryDistributed IntelligenceML for CommunicationIntegrated Sensing and Communication