High-Probability Convergence in Decentralized Stochastic Optimization with Gradient Tracking

📅 2026-04-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of high-probability convergence guarantees in decentralized stochastic optimization, particularly under data heterogeneity and non-strongly convex settings where existing methods rely on overly restrictive assumptions. The paper proposes a gradient-tracking-based decentralized stochastic gradient descent algorithm (GT-DSGD) that, under mild sub-Gaussian noise conditions, establishes the first high-probability convergence guarantee for a bias-corrected decentralized method, thereby bridging the theoretical gap between high-probability and mean-square-error analyses. The analysis shows that GT-DSGD achieves optimal high-probability convergence rates of $O(\log(1/\delta)/\sqrt{nT})$ for non-convex objectives and $O(\log(1/\delta)/(nT))$ under the Polyak–Łojasiewicz condition, with experiments demonstrating its clear superiority over current state-of-the-art approaches.
📝 Abstract
We study high-probability (HP) convergence guarantees in decentralized stochastic optimization, where multiple agents collaborate to jointly train a model over a network. Existing HP results in decentralized settings almost exclusively focus on the Decentralized Stochastic Gradient Descent ($\mathtt{DSGD}$) algorithm, which requires strong assumptions, such as bounded data heterogeneity, or strong convexity of each agent's cost. This is contrary to the mean-squared error (MSE) results, where methods incorporating bias-correction techniques are known to converge under relaxed assumptions and achieve better practical performance. In this paper we provide the first step toward bridging the gap, by studying HP convergence of $\mathtt{DSGD}$ incorporating the gradient tracking technique, in the presence of noise satisfying a relaxed sub-Gaussian condition. We show that the resulting method, dubbed $\mathtt{GT-DSGD}$, achieves order-optimal HP convergence rates for both non-convex and Polyak-Łojasiewicz costs, of order $\mathcal{O}\Big(\frac{\log(1/δ)}{\sqrt{nT}}\Big)$ and $\mathcal{O}\Big(\frac{\log(1/δ)}{nT}\Big)$, respectively, where $n$ is the number of agents, $T$ is the time horizon and $δ\in (0,1)$ is the confidence parameter. Our results establish that $\mathtt{GT-DSGD}$ converges in the HP sense under the same conditions on the cost as in the MSE sense, while achieving comparable transient times. To the best of our knowledge, these are the first HP guarantees for decentralized optimization methods incorporating bias-correction. Numerical experiments on real and synthetic data verify our theoretical findings, underlining the superior performance of $\mathtt{GT-DSGD}$ and highlighting that the benefits of incorporating bias-correction are also maintained in the HP sense.
Problem

Research questions and friction points this paper is trying to address.

decentralized stochastic optimization
high-probability convergence
gradient tracking
bias-correction
non-convex optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

gradient tracking
high-probability convergence
decentralized stochastic optimization
bias-correction
Polyak-Łojasiewicz condition
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Aleksandar Armacki
École Polytechnique Fédérale de Lausanne, Lausanne, Switzerland
H
Haoyuan Cai
École Polytechnique Fédérale de Lausanne, Lausanne, Switzerland
A
Ali H. Sayed
École Polytechnique Fédérale de Lausanne, Lausanne, Switzerland