minimax lower bounds

Deriving information-theoretic lower bounds on estimation error or policy regret and characterizing fundamental impossibility results. This involves constructing adversarial instances and reductions to show unavoidable tradeoffs and then comparing or matching algorithmic performance to those limits.

minimaxlowerbounds

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Regret Bounds for Robust Online Decision Making

Apr 09, 2025
AA
Alexander Appel
🏛️ Technion - Israel Institute of Technology

This paper addresses robust online decision-making under structured observations. We propose a novel framework based on multi-distribution probabilistic models: for each action, the outcome distribution belongs to a convex set, and the environment may adversarially select the true distribution from this set in a non-stationary, history-dependent manner. This setting relaxes the strong realizability assumptions prevalent in classical reinforcement learning and multi-armed bandits, enhancing practical relevance. We establish, for the first time, a theoretical foundation for robust multi-distribution modeling and introduce the notion of “power-law learnability,” providing its complete characterization. Building upon this, we derive tight regret upper and lower bounds for robust linear bandits and tabular robust online reinforcement learning—substantially improving upon prior results. Our analysis identifies power-law learnability as the fundamental scaling law governing learnability in this framework.

Derives regret bounds for adversarial nonoblivious environmentsGeneralizes decision making with robust multivalued modelsImproves state-of-the-art in robust linear bandits and RL

On Bits and Bandits: Quantifying the Regret-Information Trade-off

May 26, 2024
IS
Itai Shufaro
🏛️ Technion | ENSAE | Nvidia Research

This work investigates the quantitative trade-off between accumulated information (in bits) and cumulative regret (i.e., reward loss) in sequential decision-making. Under the Bayesian setting, we establish the first information-dependent lower bound on regret. Methodologically, we unify information theory, Bayesian optimal decision analysis, and regret decomposition to construct a coherent theoretical framework wherein information and regret are exchangeable—systematically integrating and generalizing several classical regret lower bounds. We empirically validate the tightness of our bounds via question-answering tasks using large language models (LLMs). Experimental results demonstrate that incorporating information-aware policies significantly reduces regret in LLM-based QA, providing empirical evidence for the validity and practical utility of a quantifiable information–regret trade-off.

Bayesian regret boundsRegret-information trade-offSequential decision problems

Beyond Worst-Case Online Classification: VC-Based Regret Bounds for Relaxed Benchmarks

Apr 14, 2025
OM
Omar Montasser
🏛️ Yale University | MIT | UC Berkeley

Classical regret analysis in online binary classification targets only the minimal 0–1 loss, yielding loose bounds dependent on the Littlestone dimension and failing to capture robust generalization. Method: We replace the standard benchmark with robust relaxed benchmarks—such as adversarial robustness to input perturbations, Gaussian-smoothed performance, and margin guarantees—and introduce, for the first time, VC dimension and metric entropy of the instance space into relaxed regret analysis, enabling margin-aware online algorithms and a novel regret decomposition framework. Contributions/Results: We derive tight upper bounds depending solely on VC dimension and metric entropy; the dependence on generalized margin γ improves from polynomial in prior work to optimal O(log(1/γ)), matched by a corresponding lower bound. Our framework unifies adversarial robustness, smooth learning, and VC theory, substantially enhancing both the practical applicability and theoretical tightness of generalization guarantees.

Achieves VC-based regret bounds with logarithmic margin dependenceMeasures regret against predictors robust to input perturbationsShifts focus from worst-case binary loss to relaxed benchmarks

This study addresses the noiseless inverse optimization problem, aiming to accurately infer the parameters of a decision-maker’s objective function from observed context–action pairs and to establish theoretical guarantees on its generalization performance. By integrating statistical learning theory with an analogy to best-arm identification, the work derives, for the first time, a tight high-probability generalization upper bound of $O(d/T)$ and proves that this rate is a fundamental lower bound for all consistent estimators. These findings reveal that, under this setting, stochastic inverse optimization inherently exhibits adversarial characteristics. The paper further proposes a parameter-free, computationally efficient algorithm and validates both the theoretical bounds and the predicted convergence rate through empirical experiments.

generalization boundsinverse optimizationnoiseless setting

Data-Dependent Regret Bounds for Constrained MABs

May 26, 2025
GG
Gianmarco Genalti
🏛️ Politecnico di Milano

This paper establishes the first data-dependent regret bound for constrained multi-armed bandits (MAB). We consider an adversarial loss setting with stochastic hard constraints. To address this, we propose a novel algorithm integrating online mirror descent, constraint drift compensation, and adaptive confidence intervals. We rigorously prove that the dynamic regret decomposes into two fundamental terms: “constraint-satisfaction hardness” and “unconstrained learning complexity,” and derive a matching information-theoretic lower bound. Our upper bound is tight—matching this lower bound—and constitutes the first provably optimal data-dependent regret bound for constrained MAB. Notably, when constraints are satisfied with high probability, our bound significantly improves upon the classical $widetilde{mathcal{O}}(sqrt{T})$ guarantee. Furthermore, we extend our framework to soft constraints and introduce new analytical tools for handling stochastic constraint violations.

Derive regret bounds with adversarial losses and stochastic constraintsDesign algorithm ensuring hard constraints with high probabilityStudy data-dependent regret bounds in constrained MABs

Latest Papers

What's happening recently
View more

This work addresses optimization problems reliant on machine learning predictions when reliable error bounds are unavailable, a setting where traditional robust and regret-based approaches struggle to deliver effective performance guarantees. The paper introduces the Global Adversarial Regret Optimization (GARO) framework, which generalizes the notion of adversarial regret globally and provides unified absolute or relative performance guarantees for uncertainties of arbitrary magnitude—without requiring probabilistic calibration of uncertainty sets. By extending Lepski’s adaptive method to downstream decision-making and leveraging affine worst-case cost functions with polyhedral norm-based uncertainty sets, GARO is exactly reformulated into a tractable optimization problem, accompanied by a constraint generation algorithm with convergence guarantees. Empirical results demonstrate that GARO achieves a superior trade-off between worst-case and average out-of-sample performance while offering stronger global assurances.

adversarial regretdecision-making under uncertaintyrobust optimization

This work establishes high-probability regret bounds for empirical risk minimization (ERM) and extends them to learning problems involving nuisance components, such as causal inference, missing data, and domain adaptation. By employing a three-step approach—elementary inequality, localized uniform concentration bounds, and a fixed-point argument—combined with a key radius defined via local Rademacher complexity, the study characterizes convergence rates in a modular analytical framework. This framework unifies the treatment of standard and nuisance-augmented ERM, explicitly decomposing statistical and approximation errors, and provides sufficient conditions for fast convergence. It recovers classical rates for VC-subgraph classes, Sobolev/Hölder spaces, and bounded variation function classes, and delivers transferable regret guarantees for orthogonal learning settings.

Empirical Risk MinimizationLocalized Rademacher ComplexityNuisance Components

This work addresses adversarial bandit optimization under a global perturbation budget, where in each round the loss consists of a linear function plus an action-dependent perturbation term, with the total perturbation constrained globally. Within this non-convex and non-smooth setting, the paper establishes—for the first time—both expected and high-probability regret upper bounds under a global perturbation budget, improving upon the classical high-probability regret bound in the unperturbed case. Additionally, it provides a matching lower bound on the expected regret. The analysis combines techniques from adversarial bandits, perturbation modeling, and refined probabilistic arguments, offering rigorous theoretical guarantees for online decision-making in perturbed environments.

adversarial banditglobally bounded perturbationslinear losses

This work investigates the fundamental performance limits of learning and estimation tasks within an information-theoretic framework, independent of the computational capabilities of specific algorithms. By integrating tools from information theory and statistical learning theory—including metric entropy, VC dimension, Rademacher complexity, mutual information, and relative entropy—it systematically derives multiple upper bounds on generalization error. Simultaneously, leveraging Fano’s inequality together with covering and packing numbers, the study establishes information-theoretic lower bounds on minimax risk. The analysis unifies two complementary paradigms: one grounded in the geometric structure of metric spaces and the other based on information-theoretic measures. This synthesis yields a rigorous and broadly applicable theoretical framework for characterizing the optimal performance boundaries inherent to learning and estimation problems.

estimationgeneralization errorinformation-theoretic limits

This work addresses online convex optimization with strongly convex losses under three settings: full information, bandit feedback with noise, and stochastic constraints. It establishes high-probability regret bounds that depend only on the noise level σ, rather than the gradient norm bound G. By introducing an exponential supermartingale analysis, the approach circumvents the bounded-difference assumption inherent in Freedman’s inequality. The analysis reveals a linear dependence on log(1/δ) in the bandit setting—distinct from the full-information case—and, for the first time, provides simultaneous high-probability guarantees on both regret and long-term constraint violation under stochastic constraints. The main theoretical results include an O(σ√T log(1/δ)) regret bound in the full-information setting, an Ω(log(1/δ)) lower bound on the confidence cost in the bandit setting, and, under stochastic constraints, O(√(T log(m/δ))) regret with constraint violation bounded by O(√T/(ζδ) + m√(T log(m/δ))). Empirical experiments corroborate these theoretical predictions.

Bandit FeedbackHigh-Probability Regret BoundsNoise Adaptivity

Hot Scholars

MH

Min-hwan Oh

Seoul National University
Reinforcement LearningBandit AlgorithmsMachine Learning
EL

Euiwoong Lee

University of Michigan
Theoretical computer scienceApproximation algorithmsHardness of approximation
BH

Bernhard Haeupler

INSAIT (University of Sofia, Bulgaria) & ETH Zurich
Theoretical Computer Science
PM

Pasin Manurangsi

Google Research
Theoretical Computer ScienceDifferential PrivacyApproximation AlgorithmsHardness of Approximation