Score
Designs, implements, and analyzes follow-the-regularized-leader (FTRL) style online learning algorithms for decentralized networks that receive only bandit (partial) feedback, including procedures to estimate gradients from bandit observations and to incorporate those estimates into decentralized OCO updates. This work also builds and studies communication-compressed consensus and D-OCO protocols that integrate FTRL updates, with formal analyses of regret guarantees and communication–computation tradeoffs.
This work addresses the challenge of simultaneously achieving tight regret bounds and high communication efficiency in decentralized online convex optimization under compressed communication. For the first time, it introduces the Follow-the-Regularized-Leader (FTRL) framework into this setting, leveraging its dual-space update mechanism to naturally integrate compressed averaging consensus protocols. Building upon this insight, the paper proposes novel algorithms for both full-information and bandit feedback scenarios. The resulting methods offer notably simpler algorithmic design and theoretical analysis: in the full-information case, they attain the best-known regret bound; under bandit feedback, they significantly improve both regret and communication complexity, outperforming existing approaches.
Existing adaptive learning rate methods are scarce for online learning problems with minimax regret of order Θ(T²⁄₃), such as partial monitoring, graph bandits, and costly-observation multi-armed bandits. Method: We propose the first simple, adaptive learning rate framework tailored to this regret scale, built upon Follow-the-Regularized-Leader (FTRL) with Tsallis entropy regularization. Our design achieves unified optimization in both stochastic and adversarial environments by precisely balancing stability, penalty, and bias terms. Contribution/Results: Unlike existing Best-of-Both-Worlds algorithms relying on intricate rate constructions, our approach significantly simplifies algorithm design while attaining tighter regret upper bounds across all three Θ(T²⁄₃) problem classes. It is the first method to simultaneously improve performance in both stochastic and adversarial settings. The resulting learning rate depends only on logarithmic-scale terms, ensuring both theoretical tightness and practical applicability.
This work addresses the design of parameter-free algorithms with provable network regret guarantees for decentralized online learning. We propose the first algorithmic family that integrates multi-agent coin-flipping mechanisms with gossip communication, introducing a novel “betting function” analytical framework to uniformly characterize both individual and network-level regret behavior—thereby significantly simplifying multi-agent decentralized regret analysis. Theoretically, our method achieves a sublinear network regret bound of $O(sqrt{T})$ over connected communication graphs, without requiring any hyperparameter tuning. Empirical evaluation on synthetic benchmarks and real-world distributed sensing tasks confirms its robustness and efficiency. Key contributions include: (i) the first incorporation of coin-flipping strategies into decentralized online learning; (ii) the establishment of a general, parameter-free analytical paradigm; and (iii) a scalable, communication-efficient framework for distributed collaborative learning.
This work addresses the challenge of unknown and time-varying feedback delays in decentralized online convex optimization, where existing methods rely on prior knowledge of the total delay and yield suboptimal regret bounds. The authors propose a novel decentralized online learning algorithm that integrates adaptive learning rates with a gossip-based communication protocol, enabling agents to locally estimate delays and collaboratively optimize without requiring any prior information on the total delay. For the first time, the method achieves tight regret bounds under such agnostic conditions: $O(N\sqrt{d_{\text{tot}}} + N\sqrt{T}/(1-\sigma^2)^{1/4})$ for general convex functions and $O(N\delta_{\max} \ln T / \alpha)$ for strongly convex functions. The theoretical analysis leverages spectral graph theory, and empirical evaluations demonstrate superior performance over existing baselines.
This paper addresses the loose dynamic regret bounds of Follow-the-Regularized-Leader (FTRL) in dynamic online convex optimization (OCO), identifying the root cause as the decoupling between state updates and iterates—not the projection mechanism, as conventionally assumed. To overcome this, we propose a novel analytical framework integrating optimistic prediction of future costs with linearized gradient pruning over historical gradients. Our approach employs recursive regularization to tightly couple states and iterates, enabling loop-free optimistic design and continuous interpolation between greediness and agility. The framework recovers classical dynamic regret upper bounds as special cases, yields finer-grained control over regret terms, and achieves the optimal $O(sqrt{T})$ dynamic regret over compact domains—without increasing gradient queries or memory overhead.
This work addresses the challenge of error accumulation caused by communication compression in decentralized online convex optimization. The authors propose DECO-EF, the first parameter-free compressed learning algorithm that requires no prior knowledge of the learning rate, time horizon, or comparator norm. By integrating a coin-betting mechanism, error feedback, and a differential gossip protocol, each agent maintains a clean cumulative state and a compressed tracker, transmitting only the difference between states during communication. Theoretical analysis demonstrates that DECO-EF achieves a comparator-adaptive sublinear network regret bound under compressed communication, establishing it as the first decentralized online learning algorithm with such a guarantee.
Existing computationally efficient Follow-the-Perturbed-Leader (FTPL) algorithms struggle to design adaptive learning rates that depend on the arm-selection probabilities, limiting their best-of-both-worlds (BOBW) performance across various bandit settings. This work proposes a surrogate probability function that relies solely on observable quantities, thereby introducing—for the first time within the FTPL framework—an adaptive learning rate mechanism that avoids explicit computation of true probabilities. By integrating Pareto perturbations with arbitrary shape parameter α > 1 and ideas from online learning, the method preserves computational simplicity while extending BOBW theoretical guarantees to general perturbation distributions and bandit problems with expert advice. This significantly broadens the applicability and performance boundaries of FTPL algorithms.
This work addresses the challenge of sampling violations arising from continuous relaxation and rounding in distributed online submodular maximization under partition matroid constraints across multi-agent systems. The authors propose a unified algorithmic framework applicable to both full-information and bandit feedback settings. Its key innovation lies in a bounded stochastic pipage rounding scheme, for which they establish— for the first time—that the probability of sampling violations asymptotically vanishes and the cumulative violation remains sublinear. Under both feedback models, the method achieves tight sublinear $(1-1/e)$-regret bounds, matching the performance of centralized optimal algorithms. These theoretical guarantees are corroborated by numerical experiments.
This work addresses the challenge of integrating offline data with online learning in stochastic linear bandits by proposing a novel algorithm that initially leverages offline data to form a prior and subsequently enhances online exploration in an adaptive manner, dynamically balancing the contributions of both sources. Within the structured linear bandit setting, the method is the first to simultaneously outperform purely online and purely offline strategies, achieving a sublinear regret bound with respect to the optimal action. Notably, the regret decreases as the amount of offline data increases. Theoretical analysis establishes a rigorous upper bound on regret, and extensive experiments demonstrate that the proposed approach significantly outperforms existing baselines across various settings.