decentralized ftrl with compression

Design, implement, and analyze decentralized online convex optimization algorithms based on the Follow-The-Regularized-Leader (FTRL) framework that incorporate communication compression and compressed consensus mechanisms across networked nodes. Establish regret guarantees and convergence analyses that account for compression error and network topology while quantifying per-round communication cost and trade-offs between compression and accuracy.

decentralizedftrlwithcompression

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of simultaneously achieving tight regret bounds and high communication efficiency in decentralized online convex optimization under compressed communication. For the first time, it introduces the Follow-the-Regularized-Leader (FTRL) framework into this setting, leveraging its dual-space update mechanism to naturally integrate compressed averaging consensus protocols. Building upon this insight, the paper proposes novel algorithms for both full-information and bandit feedback scenarios. The resulting methods offer notably simpler algorithmic design and theoretical analysis: in the full-information case, they attain the best-known regret bound; under bandit feedback, they significantly improve both regret and communication complexity, outperforming existing approaches.

Bandit SettingCompressed CommunicationDecentralized Online Convex Optimization

This work proposes a novel architecture based on adaptive feature fusion and dynamic reasoning to address the limited generalization of existing methods in complex scenarios. By incorporating a multi-scale context-aware module and a learnable strategy for selecting inference paths, the approach significantly enhances model robustness to out-of-distribution data. Experimental results demonstrate that the proposed method consistently outperforms current state-of-the-art models across multiple benchmark datasets, achieving an average accuracy improvement of 3.2% while maintaining low computational overhead. The primary contribution lies in the design of a general and efficient dynamic inference framework, offering a new perspective for improving the adaptability of AI systems in open-world environments.

Bandit FeedbackCompressed CommunicationConsensus

This work addresses the high communication overhead in distributed online convex optimization under large-scale streaming data by studying optimization methods with compressed communication. By integrating the Follow-the-Regularized-Leader framework, an error-feedback mechanism, and a bidirectional online compression strategy—augmented with an online-to-batch conversion—the proposed approach effectively decouples compression errors from projection errors and controls their accumulation. The paper establishes, for the first time, theoretical lower bounds on regret for both convex and strongly convex loss functions under compressed communication, and introduces an optimal algorithm that achieves these bounds. It also provides the first convergence guarantees for distributed nonsmooth optimization with compressed communication and domain constraints. The method attains optimal regret bounds of $O(\delta^{-1/2}\sqrt{T})$ and $O(\delta^{-1}\log T)$ in the convex and strongly convex settings, respectively, yielding corresponding offline convergence rates of $O(\delta^{-1/2}T^{-1/2})$ and $O(\delta^{-1}T^{-1})$.

Communication CostCompressed CommunicationDistributed Online Convex Optimization

Lower Bounds and Accelerated Algorithms in Distributed Stochastic Optimization with Communication Compression

May 12, 2023
YH
Yutong He
🏛️ Peking University | University of Pennsylvania | MetaCarbon Inc. | DAMO Academy | Alibaba Group

This work investigates the fundamental performance limits of distributed stochastic optimization under communication compression. We establish the first set of tight convergence lower bounds—covering six distinct settings formed by combining strongly convex, convex, and nonconvex objective functions with unbiased and contractive compressors. Building upon these bounds, we propose NEOLITHIC, the first compression-based algorithm achieving near-optimal rates (up to logarithmic factors) across all six settings. NEOLITHIC integrates variance reduction, momentum acceleration, and explicit compressor modeling. We prove theoretically that it matches the derived lower bounds under mild assumptions. Empirical evaluations demonstrate that, in multi-node training, NEOLITHIC reduces communication overhead by 3–5× compared to state-of-the-art compressed methods, while significantly improving convergence efficiency.

Determine performance limits of distributed stochastic optimization with communication compression.Establish lower bounds for convergence rates in various optimization settings.Propose NEOLITHIC algorithm to achieve nearly optimal convergence rates.

Decentralized Sparse Linear Regression via Gradient-Tracking: Linear Convergence and Statistical Guarantees

Jan 21, 2022
YS
Ying Sun
🏛️ The Pennsylvania State University | Texas A&M University | Purdue University | The University of California, Los Angeles

This paper studies high-dimensional sparse linear regression in a decentralized multi-agent network without a central server: each node holds only local observations, and the ambient dimension $d$ may vastly exceed the total sample size $N$. We propose a distributed projected gradient tracking algorithm, the first to achieve linear convergence rate in the decentralized setting while attaining the centralized statistical optimal error bound $O(s log d / N)$. Our analysis reveals an intrinsic coupling among network connectivity, statistical efficiency, and convergence rate. Under the condition $s log d / N = o(1)$, the algorithm achieves an $varepsilon$-optimal solution with computational complexity matching that of centralized methods; moreover, the required number of communication rounds decreases as the spectral gap of the mixing matrix increases. The framework simultaneously guarantees statistical consistency and linear convergence—thereby extending the theoretical foundations of high-dimensional decentralized learning.

Decentralized EnvironmentHigh-Dimensional DataSparse Linear Regression

Latest Papers

What's happening recently
View more

This work addresses decentralized non-smooth non-convex optimization under communication constraints by proposing a unified algorithmic framework that integrates unbiased or contractive compression (with error compensation), stochastic subgradient updates, gradient-tracking momentum, and sign-based regularization. For the first time, global convergence is rigorously established under conditions where the objective function is non-smooth and fails to satisfy Clarke regularity, leveraging differential inclusion theory to analyze consensus errors and averaged trajectories. Both theoretical analysis and empirical experiments demonstrate that the proposed method significantly reduces communication overhead while effectively preserving optimization accuracy, thereby achieving a favorable trade-off between communication efficiency and convergence performance.

communication compressiondecentralized optimizationnonsmooth nonconvex

This work addresses curvature-adaptive online optimization for non-convex loss functions, aiming to unify sublinear regret in general non-convex settings with logarithmic regret under strong convexity. The authors propose a novel curvature-adaptive Follow-the-Perturbed-Leader (FTPL) algorithm that, for the first time, incorporates a dynamic perturbation scale into the non-convex FTPL framework, automatically adapting to the local geometry of the loss without requiring prior knowledge of cumulative curvature. By combining time-varying perturbations, a Follow-the-Leader tuning rule, and an approximate offline optimization oracle, the method achieves an $O(\sqrt{T})$ regret bound for arbitrary Lipschitz non-convex losses. When cumulative curvature grows linearly—including the strongly convex case—and the oracle is sufficiently accurate, it further attains an $O(\log T)$ regret bound, which is shown to be information-theoretically optimal.

curvature adaptivitynon-convex lossesonline optimization

This work addresses decentralized stochastic smooth convex optimization over a fixed communication network, aiming to maximize the number of participating nodes $M$ under a total gradient sample budget $N$ while preserving the optimal statistical convergence rate of $O(1/\sqrt{N})$ achievable by centralized methods. To this end, the authors propose a novel algorithm that integrates accelerated gossip communication, mini-batch gradients, and a single-step delayed acceleration mechanism. This approach effectively controls the residual inconsistency among nodes and exhibits only logarithmic dependence on local data heterogeneity. The method achieves a significantly improved scalability bound of $M \lesssim \sqrt{\rho}\, N^{3/4}$, where $\rho$ denotes the network spectral gap, surpassing the previous best-known bound of $M \lesssim \rho \sqrt{N}$. Moreover, the authors establish the optimality of this bound for first-order methods within the linear span class.

decentralized optimizationgossip networkstatistical rate

This work addresses the challenge of error accumulation caused by communication compression in decentralized online convex optimization. The authors propose DECO-EF, the first parameter-free compressed learning algorithm that requires no prior knowledge of the learning rate, time horizon, or comparator norm. By integrating a coin-betting mechanism, error feedback, and a differential gossip protocol, each agent maintains a clean cumulative state and a compressed tracker, transmitting only the difference between states during communication. Theoretical analysis demonstrates that DECO-EF achieves a comparator-adaptive sublinear network regret bound under compressed communication, establishing it as the first decentralized online learning algorithm with such a guarantee.

compressed communicationdecentralized online learninggossip

Hot Scholars

DM

Daniel M. Roy

Research Director, Vector Institute; Prof., U. Toronto (Statistics, CS)
Machine learningTrustworthy AIMathematical StatisticsLearning Theory
ZH

Zhi-Hua Zhou

Nanjing University
Artificial IntelligenceMachine LearningData Mining
LH

Longbo Huang

Professor, IIIS, Tsinghua University, ACM Distinguished Scientist
Reinforcement Learning (RL)Deep RLMachine LearningStochastic Networks
YD

Yihan Du

Assistant Professor, SUTD ESD
Reinforcement LearningOnline LearningRepresentation Learning
YJ

Yu-Jie Zhang

RIKEN AIP
Machine LearningOnline LearningWeakly Supervised Learning