minimax rate analysis

Derive tight upper and lower bounds on the worst-case (minimax) risk or estimation error for statistical estimation and learning problems. Characterize how the convergence rate scales with sample size, loss, function-class smoothness and other problem parameters (e.g., model complexity or depth), and produce matching bounds that establish sample-complexity and estimator optimality.

minimaxrateanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.06
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Fast Rate Information-theoretic Bounds on Generalization Errors

Mar 26, 2023
XW
Xuetong Wu
🏛️ University of Melbourne

This work addresses the tightness of information-theoretic generalization error bounds with respect to sample size $n$, particularly the looseness of the individual-sample mutual information (ISMI) bound. To overcome the suboptimal $O(1/sqrt{n})$ convergence rate, we introduce, for the first time, an *excess risk assumption*, yielding a tight $O(1/n)$ fast-rate bound. Furthermore, we propose a novel generalization framework based on the $(eta,c)$-central condition, under which the mutual information term directly governs the convergence rate. We rigorously prove that this bound achieves the optimal $O(1/n)$ rate under standard assumptions. Empirical evaluation on canonical tasks—such as Gaussian mean estimation—demonstrates substantial improvements over existing information-theoretic bounds. The proposed framework thus bridges theoretical rigor with practical superiority, advancing both the tightness and applicability of information-theoretic generalization analysis.

Investigates tightness of generalization error boundsProposes new bounds using (η, c)-central conditionShows fast rate recovery under excess risk assumption

This work investigates the failure of generalization in empirical risk minimization (ERM) for stochastic convex optimization when the sample size scales linearly with the problem dimension. By constructing a specific convex learning instance, the authors provide the first proof that ERM can be uniquely and necessarily overfitted under this regime, thereby resolving an open problem posed by Feldman. The analysis is further extended to approximate ERM and gradient descent algorithms. Combining tools from convex optimization theory, probabilistic constructions, and dynamical systems analysis, the study establishes a lower bound of Ω(√(ηT/m^{1.5})) on the generalization error of gradient descent, significantly narrowing the exponential gap between the previously known upper bound of O(ηT/m) and the true behavior of the algorithm.

Empirical Risk MinimizationGradient DescentOverfitting

This study investigates high-confidence non-asymptotic upper and lower bounds for the minimal risk in statistical learning, circumventing the conventional reliance on boundedness assumptions of the empirical risk function. By integrating sharp forms of Talagrand’s concentration inequality—specifically the Bousquet and Klein–Rio refinements—with transport-entropy inequalities and empirical process theory, the authors derive a lower bound independent of both the number of parameters and input dimension, under Gaussian or exponential integrability assumptions. The corresponding upper bound is characterized by the interplay between sample size and the box-counting dimension of the parameter set measured in an Orlicz norm. This work thus provides a more general and non-asymptotic theoretical framework for evaluating learning algorithm performance without resorting to asymptotic approximations.

concentration inequalitiesempirical risk principleminimal risk

This work addresses the minimax estimation of discrete probability distributions under the $\ell_\infty$ norm, aiming to characterize optimal risk bounds both in expectation and with high probability. By integrating minimax theory, high-dimensional probability analysis, and empirical process techniques with constructive proofs, the study establishes the first fully computable, data-dependent tight risk bound, thereby resolving an open problem posed by Kontorovich and Painsky. The analysis also precisely identifies the structure of extremal distributions that achieve worst-case risk. These theoretical advances not only sharpen existing $\ell_\infty$ risk bounds but also lead to estimators that demonstrate superior empirical performance compared to current methods.

distribution estimationextremal distributionhigh-probability tail bounds

Minimax learning rates for estimating binary classifiers under margin conditions

May 15, 2025
JG
Jonathan Garc'ia
🏛️ University of Vienna

This paper investigates the minimax learning rates of binary classifiers under geometric margin constraints. We develop a Lebesgue-norm analytical framework constrained by Kolmogorov entropy, integrating minimax theory with geometric margin modeling, for both noiseless and noisy settings—specifically, those with level-set-type decision boundaries. Our contributions are threefold: (i) For the noiseless case, we establish the first tight lower bound on the optimal learning rate; (ii) Under Barron regularity or Hölder continuity assumptions on the decision boundary, we achieve nearly optimal $O(n^{-1})$ convergence rates; (iii) For multiple function classes—including convex boundaries—we derive matching upper and lower bounds, covering the full spectrum of rates from $O(n^{-1})$ under strong margin to $O(n^{-1/2})$ in weaker regimes. The key innovation lies in resolving the long-standing challenge of deriving tight noiseless lower bounds under geometric margin, while unifying the impact of noise structure and functional regularity on learning rates.

Deriving minimax learning rates for bounded function classesEstablishing optimal rates for Barron and Hölder function classesEstimating binary classifiers under geometric margin conditions

Latest Papers

What's happening recently
View more

This work establishes high-probability regret bounds for empirical risk minimization (ERM) and extends them to learning problems involving nuisance components, such as causal inference, missing data, and domain adaptation. By employing a three-step approach—elementary inequality, localized uniform concentration bounds, and a fixed-point argument—combined with a key radius defined via local Rademacher complexity, the study characterizes convergence rates in a modular analytical framework. This framework unifies the treatment of standard and nuisance-augmented ERM, explicitly decomposing statistical and approximation errors, and provides sufficient conditions for fast convergence. It recovers classical rates for VC-subgraph classes, Sobolev/Hölder spaces, and bounded variation function classes, and delivers transferable regret guarantees for orthogonal learning settings.

Empirical Risk MinimizationLocalized Rademacher ComplexityNuisance Components

This study investigates the theoretical foundations of operator learning, with a focus on convergence rates and fundamental statistical limits. By integrating tools from statistical learning theory, approximation theory, and the framework of holomorphic operators, the work establishes a unified error analysis framework to systematically derive generalization error bounds for empirical risk minimization. Under generalized regularity conditions, it further establishes minimax-optimal statistical lower bounds, revealing an intrinsic trade-off between sample complexity and model approximation capacity. The analysis delineates current theoretical boundaries in operator learning and identifies several key open problems, offering new perspectives to guide future theoretical advances in the field.

convergence ratesminimax theoryoperator learning

This study establishes minimax lower bounds for the estimation of Maximum Mean Discrepancy (MMD), Hilbert–Schmidt Independence Criterion (HSIC), and Kernelized Stein Discrepancy (KSD) in general topological spaces under unbounded kernel conditions. By integrating reproducing kernel Hilbert space theory, functional analysis, and a minimax information-theoretic framework, the work rigorously proves—under mild assumptions—that the optimal convergence rate for these three classes of kernel-based discrepancy measures remains $n^{-1/2}$. This result resolves a long-standing open theoretical question and extends to the estimation of mean embeddings and centered cross-covariance operators, thereby establishing the minimax optimality of their parametric convergence rates.

HSICkernel discrepancyKSD

This work investigates the fundamental performance limits of learning and estimation tasks within an information-theoretic framework, independent of the computational capabilities of specific algorithms. By integrating tools from information theory and statistical learning theory—including metric entropy, VC dimension, Rademacher complexity, mutual information, and relative entropy—it systematically derives multiple upper bounds on generalization error. Simultaneously, leveraging Fano’s inequality together with covering and packing numbers, the study establishes information-theoretic lower bounds on minimax risk. The analysis unifies two complementary paradigms: one grounded in the geometric structure of metric spaces and the other based on information-theoretic measures. This synthesis yields a rigorous and broadly applicable theoretical framework for characterizing the optimal performance boundaries inherent to learning and estimation problems.

estimationgeneralization errorinformation-theoretic limits

Hot Scholars

EH

Edward H. Kennedy

Associate Professor of Statistics & Data Science, Carnegie Mellon University
causal inferencenonparametricsmachine learninghealth & public policy
JS

Jonathan Scarlett

Associate Professor, Computer Science and Mathematics, National University of Singapore
Information TheoryMachine LearningHigh-dimensional StatisticsGroup Testing
LL

Lizhen Lin

Department of Mathematics, The University of Maryland
Geometry & StatisticsBayesian TheoryStatistics Theory of Deep LearningGeometric Deep Learning
MG

Michael Gastpar

Professor, Ecole Polytechnique Fédérale (EPFL), Switzerland
Information TheorySignal ProcessingNeuroscience
LZ

Linjun Zhang

Associate Professor of Statistics, Rutgers University
High-Dimensional StatisticsDeep LearningDifferential PrivacyAlgorithmic Fairness