regret analysis

Designs and analyzes theoretical performance guarantees for online decision algorithms by deriving, decomposing, and tightening regret bounds and related metrics across variants such as perturbation-budgeted settings, switching costs, delayed feedback, incentive-aware learners, and amortized/recovery costs. This work builds formal upper-bound proofs, amortized-regret characterizations and decompositions, and matching minimax or problem-dependent lower bounds (via adversarial-instance constructions, packing/recurrence arguments, and impossibility/separation proofs) that establish sublinear rates, delay or dimension penalties, and tradeoffs between regret and other resources.

regretanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Regret Bounds for Robust Online Decision Making

Apr 09, 2025
AA
Alexander Appel
🏛️ Technion - Israel Institute of Technology

This paper addresses robust online decision-making under structured observations. We propose a novel framework based on multi-distribution probabilistic models: for each action, the outcome distribution belongs to a convex set, and the environment may adversarially select the true distribution from this set in a non-stationary, history-dependent manner. This setting relaxes the strong realizability assumptions prevalent in classical reinforcement learning and multi-armed bandits, enhancing practical relevance. We establish, for the first time, a theoretical foundation for robust multi-distribution modeling and introduce the notion of “power-law learnability,” providing its complete characterization. Building upon this, we derive tight regret upper and lower bounds for robust linear bandits and tabular robust online reinforcement learning—substantially improving upon prior results. Our analysis identifies power-law learnability as the fundamental scaling law governing learnability in this framework.

Derives regret bounds for adversarial nonoblivious environmentsGeneralizes decision making with robust multivalued modelsImproves state-of-the-art in robust linear bandits and RL

This work addresses online convex optimization with strongly convex losses under three settings: full information, bandit feedback with noise, and stochastic constraints. It establishes high-probability regret bounds that depend only on the noise level σ, rather than the gradient norm bound G. By introducing an exponential supermartingale analysis, the approach circumvents the bounded-difference assumption inherent in Freedman’s inequality. The analysis reveals a linear dependence on log(1/δ) in the bandit setting—distinct from the full-information case—and, for the first time, provides simultaneous high-probability guarantees on both regret and long-term constraint violation under stochastic constraints. The main theoretical results include an O(σ√T log(1/δ)) regret bound in the full-information setting, an Ω(log(1/δ)) lower bound on the confidence cost in the bandit setting, and, under stochastic constraints, O(√(T log(m/δ))) regret with constraint violation bounded by O(√T/(ζδ) + m√(T log(m/δ))). Empirical experiments corroborate these theoretical predictions.

Bandit FeedbackHigh-Probability Regret BoundsNoise Adaptivity

Decision-Theoretic Approaches in Learning-Augmented Algorithms

Jan 29, 2025
SA
Spyros Angelopoulos
🏛️ Sorbonne University | International Laboratory on Learning Systems

Evaluating learning-augmented online algorithms under uncertainty remains challenging, as conventional metrics focus narrowly on worst-case prediction errors, neglecting both prediction accuracy and risk sensitivity. Method: We propose a dual-track evaluation framework grounded in decision theory, jointly incorporating distance-based prediction error quantification (deterministic aspect) and risk-sensitive modeling (stochastic aspect). By embedding decision-theoretic loss functions into online algorithm analysis, we integrate prediction error modeling with risk-controllable optimization, designing novel learning-augmented algorithms for contract scheduling and 1-max search. Contribution/Results: Our approach achieves provable robustness to prediction errors, performance guarantees with tight bounds, and explicit risk controllability. It is the first to unify prediction accuracy, worst-case robustness, and risk preference within a single theoretical framework—establishing a systematic evaluation paradigm and design principle for learning-augmented online algorithms.

Decision TheoryMachine LearningOptimal Balance

Online Inverse Linear Optimization: Improved Regret Bound, Robustness to Suboptimality, and Toward Tight Regret Analysis

Jan 24, 2025
SS
Shinsaku Sakaue
🏛️ The University of Tokyo | RIKEN AIP | Kyoto University | Hokkaido University

This paper studies online inverse linear optimization: over $T$ rounds, dynamically estimating an agent’s hidden linear objective function from its (approximately) optimal actions on time-varying feasible sets. To address large cumulative regret in high-dimensional, long-horizon settings, we propose MetaGrad-ONS—a meta-algorithmic framework—achieving the first tight standard regret bound of $O(n log T)$, improving upon prior state-of-the-art by a factor of $Omega(n^3)$. We further design a robust variant resilient to suboptimal actions, attaining a robust regret bound of $O(n log T + sqrt{Delta_T n log T})$, where $Delta_T$ quantifies action suboptimality; we prove an $Omega(n)$ lower bound, establishing near-tightness. Notably, when dimension $n = 2$, the bound reduces to $O(1)$ constant regret—revealing that the fundamental bottleneck for tight high-dimensional bounds lies in dimensional coupling, not temporal scaling.

High-dimensional SequencesOnline LearningRegret Minimization

Decoupling Learning and Decision-Making: Breaking the O(√T) Barrier in Online Resource Allocation with First-Order Methods

Feb 11, 2024
WG
Wenzhi Gao
🏛️ Stanford University | Shanghai University of Finance and Economics | Shanghai Jiao Tong University

In online linear programming for resource allocation, classical first-order methods have long been constrained by an Ω(√T) regret lower bound. This paper introduces a novel “decoupled learning and decision-making” framework that breaks this fundamental barrier for the first time. By designing a constraint-aware, two-timescale gradient update mechanism—jointly optimizing online convex optimization and dynamic decision-making—the approach achieves an O(T^{1/3}) regret upper bound. This result substantially improves upon the standard O(√T) regret of conventional first-order methods and approaches the performance of logarithmic-regret optimal algorithms. The framework provides a new paradigm for high-dimensional online resource allocation, offering both strong theoretical guarantees and practical computational efficiency.

Online Linear ProgrammingPerformance BoundResource Allocation

Latest Papers

What's happening recently
View more

This work addresses the incompatibility between traditional static Nash equilibria and individual regret in online dynamic games. To bridge this gap, the authors introduce two new performance measures: the Static Duality Gap (SDual-Gap) and the Dynamic Saddle-Point Regret (DSP-Reg). They develop a unified theoretical framework applicable to strongly convex–strongly concave functions, min-max exponentially concave functions, and those satisfying the two-sided Polyak–Łojasiewicz condition. By reducing the problem to classical online convex optimization (OCO), they design corresponding algorithms and establish tight theoretical bounds for the proposed metrics. The framework not only unifies the analysis across diverse function classes but also demonstrates practical efficacy through applications such as two-player portfolio selection, confirming its generality and real-world relevance.

cumulative saddle pointsduality gapdynamic regret

This work addresses the efficient minimization of linear and profile swap regret in online optimization. For general convex decision sets, it introduces a novel algorithm that combines the responsive approachability framework with geometric preprocessing via John’s ellipsoid, marking the first application of responsive approachability to swap regret. The method yields computationally efficient solutions with tight regret bounds: achieving $O(d^{3/2}\sqrt{T})$ linear swap regret over general convex sets and improving to $O(d\sqrt{T})$ for centrally symmetric sets, matching the established information-theoretic lower bound of $\Omega(d\sqrt{T})$. Furthermore, the algorithm simultaneously minimizes profile swap regret to guard against strategic manipulation and extends to polynomial-dimensional swap deviation sets, thereby unifying and strengthening the theoretical foundations of equilibrium computation and online learning.

approachabilitycorrelated equilibriumnon-manipulability

Hot Scholars

MH

Min-hwan Oh

Seoul National University
Reinforcement LearningBandit AlgorithmsMachine Learning
MV

Michal Valko

Chief Models Officer @ Stealth Startup, Inria & MVA - Ex: Llama at Meta; Gemini and BYOL @ Deepmind
large language modelsreasoningfine-tuningtest-time computation
GF

Gabriele Farina

Assistant Professor of Computer Science, MIT
Computational Game TheoryOptimizationEconomics and Computation