Score
Designs and analyzes policy optimization algorithms that implement optimistic follow-the-regularized-leader (FTRL) updates for sequential decision-making; specifically, one constructs the update rules and theoretical analyses that adapt to adversarial and stochastic losses and establish first-/second-order, path-length, and gap-dependent regret or sample-complexity guarantees.
This paper addresses the multi-armed bandit problem under hybrid adversarial-stochastic reward distributions with unknown structure. We propose Fuzzy-FTPL, a novel Follow-The-Perturbed-Leader algorithm incorporating fuzzy distributional perturbations. Methodologically, we introduce the “principle of optimism under fuzziness,” enabling robust decision-making against unknown but bounded perturbations; further, we unify the theoretical elegance of Follow-The-Regularized-Leader (FTRL) with the computational efficiency of FTPL, achieving, for the first time, full coverage of mainstream optimal FTRL algorithms within a single framework. Our theoretical contributions include: (i) establishing a new paradigm of fuzzy robustness modeling, attaining the optimal regret bound $O(sqrt{KT})$; (ii) significantly improving computational efficiency—up to $10^4 imes$ faster than standard FTRL; and (iii) rigorously proving the optimality-equivalence between FTPL and FTRL, resolving a long-standing open problem in online learning.
This paper addresses the loose dynamic regret bounds of Follow-the-Regularized-Leader (FTRL) in dynamic online convex optimization (OCO), identifying the root cause as the decoupling between state updates and iterates—not the projection mechanism, as conventionally assumed. To overcome this, we propose a novel analytical framework integrating optimistic prediction of future costs with linearized gradient pruning over historical gradients. Our approach employs recursive regularization to tightly couple states and iterates, enabling loop-free optimistic design and continuous interpolation between greediness and agility. The framework recovers classical dynamic regret upper bounds as special cases, yields finer-grained control over regret terms, and achieves the optimal $O(sqrt{T})$ dynamic regret over compact domains—without increasing gradient queries or memory overhead.
This work investigates the time complexity of the Follow the Regularized Leader (FTRL) algorithm in converging to Nash equilibria in potential games. By constructing explicit instances, it establishes—for the first time—an exponential lower bound on the convergence time of FTRL in two-player potential games under any permutation-invariant regularizer, and a doubly exponential lower bound in the multi-player setting. The study further reveals that the FTRL dynamics admit a potential function structure and demonstrates their equivalence to mirror descent and fictitious play under appropriate conditions. These results imply that FTRL and its variants, such as multiplicative weights update, require exponential time to converge, whereas lazy alternating no-regret dynamics achieve an upper bound of $\exp(O(1/\varepsilon^2))$, matching the lower bounds up to exponential order.
This work addresses the open problem posed by Syrgkanis et al. concerning last-iterate convergence of Optimistic Multiplicative Weights Update (OMWU) in constrained convex-concave minimax optimization—particularly relevant to zero-sum games and GANs. We establish, for the first time, global convergence of OMWU to exact saddle points under general convex constraints, without requiring unconstrained domains or strong regularization assumptions. Our analysis introduces a novel framework combining monotonic KL-divergence descent with local contraction mapping properties, integrating fixed-point theory and contraction mapping techniques. This approach overcomes key limitations of prior analyses reliant on either unbounded domains or stringent regularization. The result provides the first rigorous theoretical guarantee for OMWU’s last-iterate convergence in constrained saddle-point optimization and furnishes a principled foundation for termination criteria in practical iterative implementations.
This paper studies how decision-makers learn online in repeated stochastic choice settings without knowledge of true option utilities, under the Random Utility Model (RUM). We embed RUM into an online decision-making framework and design gradient-based learning algorithms. We establish, for the first time, that multi-class RUM satisfies Hannan consistency. Moreover, we prove a rigorous equivalence between RUM-based learning and Follow-the-Regularized-Leader (FTRL), thereby providing a microeconomic foundation for FTRL. The framework is extended to model recency bias, no-regret learning in games, and prediction market mechanism design. Theoretically, it guarantees convergence of long-run average payoff to that of the optimal fixed strategy and satisfies no-regret properties. Empirically, the approach significantly improves behavioral prediction accuracy and mechanistic interpretability across three canonical economic domains.
This study investigates the exploitability of Follow-the-Regularized-Leader (FTRL) learners with fixed step sizes in two-player zero-sum games when facing oracle optimizers. By integrating game theory, online learning theory, and stochastic game models, and distinguishing between fixed and alternating optimizer settings, the work establishes— for the first time—that exploitability is an inherent property of the FTRL family. The core contributions include a geometric dichotomy based on the steepness of the regularizer, a sensitivity metric quantifying vulnerability to strategic manipulation, and theoretical guarantees showing a lower bound of Ω(N/η) on exploitability under fixed optimizers, as well as a high-probability surplus of Ω(ηT/poly(n,m)) in cumulative payoff under alternating optimizers.
Existing computationally efficient Follow-the-Perturbed-Leader (FTPL) algorithms struggle to design adaptive learning rates that depend on the arm-selection probabilities, limiting their best-of-both-worlds (BOBW) performance across various bandit settings. This work proposes a surrogate probability function that relies solely on observable quantities, thereby introducing—for the first time within the FTPL framework—an adaptive learning rate mechanism that avoids explicit computation of true probabilities. By integrating Pareto perturbations with arbitrary shape parameter α > 1 and ideas from online learning, the method preserves computational simplicity while extending BOBW theoretical guarantees to general perturbation distributions and bandit problems with expert advice. This significantly broadens the applicability and performance boundaries of FTPL algorithms.
This work addresses the challenge of learning in partially observable Markov games against strategy-dependent adaptive adversaries, where standard regret notions become inadequate. The authors propose an optimistic maximum-likelihood algorithm that partitions the learning horizon into geometrically increasing phases, leveraging cumulative confidence sets and the aggregate Eluder dimension of the observable operator class to achieve low policy regret. Their theoretical analysis establishes matching upper and lower bounds, demonstrating that the algorithm attains $\widetilde{O}(\sqrt{T})$ policy regret under fixed parameters and achieves minimax optimality with respect to both $\sqrt{T}$ and the aggregate Eluder dimension. The approach is further extended to handle horizon-adaptive and geometrically decaying memory adversaries, incorporating a logarithmic policy comparison cost mechanism.