Score
Deriving information-theoretic lower bounds on estimation error or policy regret and characterizing fundamental impossibility results. This involves constructing adversarial instances and reductions to show unavoidable tradeoffs and then comparing or matching algorithmic performance to those limits.
Classical learning theory, relying on uniform convergence over hypothesis spaces, struggles to explain the strong generalization performance of over-parameterized deep neural networks. This work proposes a unified framework that, for the first time, incorporates data-dependent worst-case generalization bounds into a single template inequality. The framework systematically integrates extensions of PAC-Bayes theory, geometric and topological characterizations of optimization trajectories—such as fractal dimension and α-weighted persistence sums—and information-theoretic surrogates grounded in algorithmic stability. By enabling direct comparison among diverse generalization bounds, it reveals their intrinsic connections and distinctions, and establishes a non-vacuous, tight family of upper bounds on generalization error that effectively accounts for the empirical generalization behavior of over-parameterized models.
This paper addresses robust online decision-making under structured observations. We propose a novel framework based on multi-distribution probabilistic models: for each action, the outcome distribution belongs to a convex set, and the environment may adversarially select the true distribution from this set in a non-stationary, history-dependent manner. This setting relaxes the strong realizability assumptions prevalent in classical reinforcement learning and multi-armed bandits, enhancing practical relevance. We establish, for the first time, a theoretical foundation for robust multi-distribution modeling and introduce the notion of “power-law learnability,” providing its complete characterization. Building upon this, we derive tight regret upper and lower bounds for robust linear bandits and tabular robust online reinforcement learning—substantially improving upon prior results. Our analysis identifies power-law learnability as the fundamental scaling law governing learnability in this framework.
This work investigates the quantitative trade-off between accumulated information (in bits) and cumulative regret (i.e., reward loss) in sequential decision-making. Under the Bayesian setting, we establish the first information-dependent lower bound on regret. Methodologically, we unify information theory, Bayesian optimal decision analysis, and regret decomposition to construct a coherent theoretical framework wherein information and regret are exchangeable—systematically integrating and generalizing several classical regret lower bounds. We empirically validate the tightness of our bounds via question-answering tasks using large language models (LLMs). Experimental results demonstrate that incorporating information-aware policies significantly reduces regret in LLM-based QA, providing empirical evidence for the validity and practical utility of a quantifiable information–regret trade-off.
Classical regret analysis in online binary classification targets only the minimal 0–1 loss, yielding loose bounds dependent on the Littlestone dimension and failing to capture robust generalization. Method: We replace the standard benchmark with robust relaxed benchmarks—such as adversarial robustness to input perturbations, Gaussian-smoothed performance, and margin guarantees—and introduce, for the first time, VC dimension and metric entropy of the instance space into relaxed regret analysis, enabling margin-aware online algorithms and a novel regret decomposition framework. Contributions/Results: We derive tight upper bounds depending solely on VC dimension and metric entropy; the dependence on generalized margin γ improves from polynomial in prior work to optimal O(log(1/γ)), matched by a corresponding lower bound. Our framework unifies adversarial robustness, smooth learning, and VC theory, substantially enhancing both the practical applicability and theoretical tightness of generalization guarantees.
This study addresses the noiseless inverse optimization problem, aiming to accurately infer the parameters of a decision-maker’s objective function from observed context–action pairs and to establish theoretical guarantees on its generalization performance. By integrating statistical learning theory with an analogy to best-arm identification, the work derives, for the first time, a tight high-probability generalization upper bound of $O(d/T)$ and proves that this rate is a fundamental lower bound for all consistent estimators. These findings reveal that, under this setting, stochastic inverse optimization inherently exhibits adversarial characteristics. The paper further proposes a parameter-free, computationally efficient algorithm and validates both the theoretical bounds and the predicted convergence rate through empirical experiments.
This paper establishes the first data-dependent regret bound for constrained multi-armed bandits (MAB). We consider an adversarial loss setting with stochastic hard constraints. To address this, we propose a novel algorithm integrating online mirror descent, constraint drift compensation, and adaptive confidence intervals. We rigorously prove that the dynamic regret decomposes into two fundamental terms: “constraint-satisfaction hardness” and “unconstrained learning complexity,” and derive a matching information-theoretic lower bound. Our upper bound is tight—matching this lower bound—and constitutes the first provably optimal data-dependent regret bound for constrained MAB. Notably, when constraints are satisfied with high probability, our bound significantly improves upon the classical $widetilde{mathcal{O}}(sqrt{T})$ guarantee. Furthermore, we extend our framework to soft constraints and introduce new analytical tools for handling stochastic constraint violations.
This work addresses optimization problems reliant on machine learning predictions when reliable error bounds are unavailable, a setting where traditional robust and regret-based approaches struggle to deliver effective performance guarantees. The paper introduces the Global Adversarial Regret Optimization (GARO) framework, which generalizes the notion of adversarial regret globally and provides unified absolute or relative performance guarantees for uncertainties of arbitrary magnitude—without requiring probabilistic calibration of uncertainty sets. By extending Lepski’s adaptive method to downstream decision-making and leveraging affine worst-case cost functions with polyhedral norm-based uncertainty sets, GARO is exactly reformulated into a tractable optimization problem, accompanied by a constraint generation algorithm with convergence guarantees. Empirical results demonstrate that GARO achieves a superior trade-off between worst-case and average out-of-sample performance while offering stronger global assurances.
This work establishes high-probability regret bounds for empirical risk minimization (ERM) and extends them to learning problems involving nuisance components, such as causal inference, missing data, and domain adaptation. By employing a three-step approach—elementary inequality, localized uniform concentration bounds, and a fixed-point argument—combined with a key radius defined via local Rademacher complexity, the study characterizes convergence rates in a modular analytical framework. This framework unifies the treatment of standard and nuisance-augmented ERM, explicitly decomposing statistical and approximation errors, and provides sufficient conditions for fast convergence. It recovers classical rates for VC-subgraph classes, Sobolev/Hölder spaces, and bounded variation function classes, and delivers transferable regret guarantees for orthogonal learning settings.
This work addresses adversarial bandit optimization under a global perturbation budget, where in each round the loss consists of a linear function plus an action-dependent perturbation term, with the total perturbation constrained globally. Within this non-convex and non-smooth setting, the paper establishes—for the first time—both expected and high-probability regret upper bounds under a global perturbation budget, improving upon the classical high-probability regret bound in the unperturbed case. Additionally, it provides a matching lower bound on the expected regret. The analysis combines techniques from adversarial bandits, perturbation modeling, and refined probabilistic arguments, offering rigorous theoretical guarantees for online decision-making in perturbed environments.
This work investigates the fundamental performance limits of learning and estimation tasks within an information-theoretic framework, independent of the computational capabilities of specific algorithms. By integrating tools from information theory and statistical learning theory—including metric entropy, VC dimension, Rademacher complexity, mutual information, and relative entropy—it systematically derives multiple upper bounds on generalization error. Simultaneously, leveraging Fano’s inequality together with covering and packing numbers, the study establishes information-theoretic lower bounds on minimax risk. The analysis unifies two complementary paradigms: one grounded in the geometric structure of metric spaces and the other based on information-theoretic measures. This synthesis yields a rigorous and broadly applicable theoretical framework for characterizing the optimal performance boundaries inherent to learning and estimation problems.
This work addresses online convex optimization with strongly convex losses under three settings: full information, bandit feedback with noise, and stochastic constraints. It establishes high-probability regret bounds that depend only on the noise level σ, rather than the gradient norm bound G. By introducing an exponential supermartingale analysis, the approach circumvents the bounded-difference assumption inherent in Freedman’s inequality. The analysis reveals a linear dependence on log(1/δ) in the bandit setting—distinct from the full-information case—and, for the first time, provides simultaneous high-probability guarantees on both regret and long-term constraint violation under stochastic constraints. The main theoretical results include an O(σ√T log(1/δ)) regret bound in the full-information setting, an Ω(log(1/δ)) lower bound on the confidence cost in the bandit setting, and, under stochastic constraints, O(√(T log(m/δ))) regret with constraint violation bounded by O(√T/(ζδ) + m√(T log(m/δ))). Empirical experiments corroborate these theoretical predictions.