Score
Designs and analyzes theoretical performance guarantees for online decision algorithms by deriving, decomposing, and tightening regret bounds and related metrics across variants such as perturbation-budgeted settings, switching costs, delayed feedback, incentive-aware learners, and amortized/recovery costs. This work builds formal upper-bound proofs, amortized-regret characterizations and decompositions, and matching minimax or problem-dependent lower bounds (via adversarial-instance constructions, packing/recurrence arguments, and impossibility/separation proofs) that establish sublinear rates, delay or dimension penalties, and tradeoffs between regret and other resources.
This paper addresses robust online decision-making under structured observations. We propose a novel framework based on multi-distribution probabilistic models: for each action, the outcome distribution belongs to a convex set, and the environment may adversarially select the true distribution from this set in a non-stationary, history-dependent manner. This setting relaxes the strong realizability assumptions prevalent in classical reinforcement learning and multi-armed bandits, enhancing practical relevance. We establish, for the first time, a theoretical foundation for robust multi-distribution modeling and introduce the notion of “power-law learnability,” providing its complete characterization. Building upon this, we derive tight regret upper and lower bounds for robust linear bandits and tabular robust online reinforcement learning—substantially improving upon prior results. Our analysis identifies power-law learnability as the fundamental scaling law governing learnability in this framework.
This work addresses online convex optimization with strongly convex losses under three settings: full information, bandit feedback with noise, and stochastic constraints. It establishes high-probability regret bounds that depend only on the noise level σ, rather than the gradient norm bound G. By introducing an exponential supermartingale analysis, the approach circumvents the bounded-difference assumption inherent in Freedman’s inequality. The analysis reveals a linear dependence on log(1/δ) in the bandit setting—distinct from the full-information case—and, for the first time, provides simultaneous high-probability guarantees on both regret and long-term constraint violation under stochastic constraints. The main theoretical results include an O(σ√T log(1/δ)) regret bound in the full-information setting, an Ω(log(1/δ)) lower bound on the confidence cost in the bandit setting, and, under stochastic constraints, O(√(T log(m/δ))) regret with constraint violation bounded by O(√T/(ζδ) + m√(T log(m/δ))). Empirical experiments corroborate these theoretical predictions.
Evaluating learning-augmented online algorithms under uncertainty remains challenging, as conventional metrics focus narrowly on worst-case prediction errors, neglecting both prediction accuracy and risk sensitivity. Method: We propose a dual-track evaluation framework grounded in decision theory, jointly incorporating distance-based prediction error quantification (deterministic aspect) and risk-sensitive modeling (stochastic aspect). By embedding decision-theoretic loss functions into online algorithm analysis, we integrate prediction error modeling with risk-controllable optimization, designing novel learning-augmented algorithms for contract scheduling and 1-max search. Contribution/Results: Our approach achieves provable robustness to prediction errors, performance guarantees with tight bounds, and explicit risk controllability. It is the first to unify prediction accuracy, worst-case robustness, and risk preference within a single theoretical framework—establishing a systematic evaluation paradigm and design principle for learning-augmented online algorithms.
This paper studies online inverse linear optimization: over $T$ rounds, dynamically estimating an agent’s hidden linear objective function from its (approximately) optimal actions on time-varying feasible sets. To address large cumulative regret in high-dimensional, long-horizon settings, we propose MetaGrad-ONS—a meta-algorithmic framework—achieving the first tight standard regret bound of $O(n log T)$, improving upon prior state-of-the-art by a factor of $Omega(n^3)$. We further design a robust variant resilient to suboptimal actions, attaining a robust regret bound of $O(n log T + sqrt{Delta_T n log T})$, where $Delta_T$ quantifies action suboptimality; we prove an $Omega(n)$ lower bound, establishing near-tightness. Notably, when dimension $n = 2$, the bound reduces to $O(1)$ constant regret—revealing that the fundamental bottleneck for tight high-dimensional bounds lies in dimensional coupling, not temporal scaling.
In online linear programming for resource allocation, classical first-order methods have long been constrained by an Ω(√T) regret lower bound. This paper introduces a novel “decoupled learning and decision-making” framework that breaks this fundamental barrier for the first time. By designing a constraint-aware, two-timescale gradient update mechanism—jointly optimizing online convex optimization and dynamic decision-making—the approach achieves an O(T^{1/3}) regret upper bound. This result substantially improves upon the standard O(√T) regret of conventional first-order methods and approaches the performance of logarithmic-regret optimal algorithms. The framework provides a new paradigm for high-dimensional online resource allocation, offering both strong theoretical guarantees and practical computational efficiency.
This work addresses the incompatibility between traditional static Nash equilibria and individual regret in online dynamic games. To bridge this gap, the authors introduce two new performance measures: the Static Duality Gap (SDual-Gap) and the Dynamic Saddle-Point Regret (DSP-Reg). They develop a unified theoretical framework applicable to strongly convex–strongly concave functions, min-max exponentially concave functions, and those satisfying the two-sided Polyak–Łojasiewicz condition. By reducing the problem to classical online convex optimization (OCO), they design corresponding algorithms and establish tight theoretical bounds for the proposed metrics. The framework not only unifies the analysis across diverse function classes but also demonstrates practical efficacy through applications such as two-player portfolio selection, confirming its generality and real-world relevance.
This work addresses the efficient minimization of linear and profile swap regret in online optimization. For general convex decision sets, it introduces a novel algorithm that combines the responsive approachability framework with geometric preprocessing via John’s ellipsoid, marking the first application of responsive approachability to swap regret. The method yields computationally efficient solutions with tight regret bounds: achieving $O(d^{3/2}\sqrt{T})$ linear swap regret over general convex sets and improving to $O(d\sqrt{T})$ for centrally symmetric sets, matching the established information-theoretic lower bound of $\Omega(d\sqrt{T})$. Furthermore, the algorithm simultaneously minimizes profile swap regret to guard against strategic manipulation and extends to polynomial-dimensional swap deviation sets, thereby unifying and strengthening the theoretical foundations of equilibrium computation and online learning.