nonconvex regret minimization

Designs and analyzes online algorithms and update rules that minimize cumulative regret against sequences of nonconvex loss functions by constructing and optimizing convex linearized surrogate losses or by invoking optimization-oracle–based updates each round. Builds surrogate linearization methods, optimization-oracle updates, and the accompanying theoretical analyses that establish conditions (e.g., regularity or approximation requirements) under which sublinear or small rp-regret is achieved.

nonconvexregretminimization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.13
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work proposes a black-box framework that unifies stochastic non-convex optimization by reducing it to static regret minimization in online convex optimization. By introducing a predictable gradient tracker and leveraging a black-box online learner to adaptively select a preconditioner for generating update directions, the method relies solely on standard static regret guarantees to handle both smooth and non-smooth non-convex objectives. This approach provides a theoretical justification for adaptive algorithms such as AdaGrad, resolving an open problem posed by Chen and Hazan (2024). In terms of convergence, the method achieves the classical $O(1/\sqrt{T})$ rate for smooth objectives and attains the optimal $O(T^{-2/7})$ rate for converging to a Goldstein stationary point under non-smooth Lipschitz non-convex settings.

black-box reductionnonconvex optimizationonline convex optimization

Online Inverse Linear Optimization: Improved Regret Bound, Robustness to Suboptimality, and Toward Tight Regret Analysis

Jan 24, 2025
SS
Shinsaku Sakaue
🏛️ The University of Tokyo | RIKEN AIP | Kyoto University | Hokkaido University

This paper studies online inverse linear optimization: over $T$ rounds, dynamically estimating an agent’s hidden linear objective function from its (approximately) optimal actions on time-varying feasible sets. To address large cumulative regret in high-dimensional, long-horizon settings, we propose MetaGrad-ONS—a meta-algorithmic framework—achieving the first tight standard regret bound of $O(n log T)$, improving upon prior state-of-the-art by a factor of $Omega(n^3)$. We further design a robust variant resilient to suboptimal actions, attaining a robust regret bound of $O(n log T + sqrt{Delta_T n log T})$, where $Delta_T$ quantifies action suboptimality; we prove an $Omega(n)$ lower bound, establishing near-tightness. Notably, when dimension $n = 2$, the bound reduces to $O(1)$ constant regret—revealing that the fundamental bottleneck for tight high-dimensional bounds lies in dimensional coupling, not temporal scaling.

High-dimensional SequencesOnline LearningRegret Minimization

This work addresses the problem of maximizing online non-monotone DR-submodular functions over downward-closed convex sets, where existing projection-free methods suffer from suboptimal regret bounds and restrictive feedback assumptions. By introducing exponential reparameterization, scaling parameters, and a surrogate potential function, the authors present the first reduction of this problem to online linear optimization, establishing a $1/e$-linearization framework. Their algorithm achieves $O(T^{1/2})$ static regret with only a single gradient query per round. The approach unifies support for multiple feedback models—including semi-bandit, full-information, and zeroth-order settings—and provides both adaptive and dynamic regret guarantees. The resulting performance strictly improves upon the best-known results in the literature.

down-closed convex setsDR-submodular maximizationnon-monotone

A Modern Introduction to Online Learning

Dec 31, 2019
FO
Francesco Orabona
🏛️ KAUST

This work addresses worst-case online optimization, unifying the study of online convex and non-convex optimization over Euclidean and non-Euclidean domains—including simplices and matrix manifolds—under a regret minimization framework. We propose a parameter-free, adaptive algorithmic framework that supports unbounded decision sets and unknown gradient magnitudes. Unifying online mirror descent (OMD) and follow-the-regularized-leader (FTRL), we reformulate first- and second-order methods and, for the first time, integrate convex surrogate losses, randomization schemes, and multi-armed bandit feedback—both adversarial and stochastic—into this coherent paradigm. Our theoretical analysis is self-contained, elementary, and accessible without prerequisites; all algorithms achieve tight, optimal regret bounds. The resulting framework establishes a universal, concise, and pedagogically transparent foundation for modern online learning, substantially lowering both theoretical barriers and practical implementation complexity.

Addresses non-convex losses using surrogate methods and randomizationIntroduces online learning via convex optimization for regret minimizationPresents adaptive algorithms for unbounded domains and parameter tuning

Decoupling Learning and Decision-Making: Breaking the O(√T) Barrier in Online Resource Allocation with First-Order Methods

Feb 11, 2024
WG
Wenzhi Gao
🏛️ Stanford University | Shanghai University of Finance and Economics | Shanghai Jiao Tong University

In online linear programming for resource allocation, classical first-order methods have long been constrained by an Ω(√T) regret lower bound. This paper introduces a novel “decoupled learning and decision-making” framework that breaks this fundamental barrier for the first time. By designing a constraint-aware, two-timescale gradient update mechanism—jointly optimizing online convex optimization and dynamic decision-making—the approach achieves an O(T^{1/3}) regret upper bound. This result substantially improves upon the standard O(√T) regret of conventional first-order methods and approaches the performance of logarithmic-regret optimal algorithms. The framework provides a new paradigm for high-dimensional online resource allocation, offering both strong theoretical guarantees and practical computational efficiency.

Online Linear ProgrammingPerformance BoundResource Allocation

Latest Papers

What's happening recently
View more

This work investigates how to leverage a sublinear number of noisy pairwise probes—comparisons indicating which of two points incurs lower loss—to improve worst-case regret in online convex optimization (OCO). The authors introduce a unified probing model and establish, for the first time, that even with only \(k = o(T)\) such noisy comparisons, the regret bound of full-feedback OCO can be significantly enhanced. By integrating a continuous exponential weights algorithm with variance-reduction analysis, they characterize the second-order effect of probing and derive an almost-tight regret upper bound of \(O\big(\min\{\sqrt{dT \ln T},\, dT \ln T / (k|1 - 2\delta|)\}\big)\), where \(\delta\) denotes the noise level. This bound is theoretically optimal in the horizon \(T\), number of probes \(k\), noise parameter \(\delta\), and number of experts \(d\) (in the finite action setting).

Noisy ProbesOnline Convex OptimizationPairwise Feedback

This work addresses online convex optimization under adversarial constraints, aiming to simultaneously minimize static regret and cumulative constraint violation. The authors propose a novel projection-based online algorithm that leverages the geometric properties of self-contracted curves, making decisions before observing the loss and constraint functions at each round. In the strongly convex setting, the algorithm achieves an $O(\log T)$ bound on both static regret and cumulative constraint violation—the latter improving upon the previous best-known rate of $O(\sqrt{T \log T})$. For general convex objectives, it attains $O(\sqrt{T})$ regret and $O(\sqrt{T})$ cumulative constraint violation, matching the optimal order in both measures.

adversarial constraintsConstrained Online Convex Optimizationcumulative constraint violation

Hot Scholars

KZ

Kaiqing Zhang

Assistant Professor, University of Maryland, College Park
Systems and ControlGame TheoryMachine LearningComputation
AO

Asuman Ozdaglar

Mathworks Professor, EECS, MIT
Optimization and Game TheoryMachine LearningEconomic and Social Networks
AD

Augustinos D. Saravanos

Postdoctoral Researcher, Massachusetts Institute of Technology
OptimizationMachine LearningControl TheoryMulti-Agent Systems
VM

Victor Magron

CNRS
Polynomial optimizationquantum informationdynamical systemsdeep learning