Score
Theoretical analysis that derives worst-case estimation or prediction error rates (minimax rates) under specified model classes and losses, often using random embeddings, augmentation schemes, or functional-analytic tools. It focuses on characterizing procedures that minimize maximum risk and proving rate-optimality for estimators under constrained scenarios.
This paper investigates the minimax learning rates of binary classifiers under geometric margin constraints. We develop a Lebesgue-norm analytical framework constrained by Kolmogorov entropy, integrating minimax theory with geometric margin modeling, for both noiseless and noisy settings—specifically, those with level-set-type decision boundaries. Our contributions are threefold: (i) For the noiseless case, we establish the first tight lower bound on the optimal learning rate; (ii) Under Barron regularity or Hölder continuity assumptions on the decision boundary, we achieve nearly optimal $O(n^{-1})$ convergence rates; (iii) For multiple function classes—including convex boundaries—we derive matching upper and lower bounds, covering the full spectrum of rates from $O(n^{-1})$ under strong margin to $O(n^{-1/2})$ in weaker regimes. The key innovation lies in resolving the long-standing challenge of deriving tight noiseless lower bounds under geometric margin, while unifying the impact of noise structure and functional regularity on learning rates.
This paper identifies an inherent suboptimality of Empirical Risk Minimization (ERM) under squared loss: its bias term dominates the estimation error, preventing attainment of the minimax optimal convergence rate; in contrast, the variance term achieves the minimax rate—a fact previously unverified rigorously. To address this, the authors establish, for the first time under random design, a non-asymptotic, sharp upper bound proving the minimax optimality of ERM’s variance term. They unify and extend Chatterjee’s admissibility theorem and the Caponnetto–Rakhlin stability result to realistic random-design settings. Furthermore, they systematically characterize the intrinsic irregularity of the empirical loss landscape for non-Donsker function classes. Integrating bias–variance decomposition, empirical process theory, and probabilistic analysis, the work delivers the most comprehensive and rigorous non-asymptotic risk decomposition for ERM to date, substantially advancing its theoretical foundations.
This paper addresses the trade-off between robustness and efficiency under model misspecification, proposing an adaptive estimation framework that does not require a pre-specified upper bound on bias. The core challenge is to construct an estimator whose worst-case risk—relative to an oracle knowing the true bias bound—is minimized. Methodologically, we formulate an adaptive shrinkage estimator via weighted convex minimax optimization, calibrated against the oracle risk, and develop a lookup-table-based fast algorithm. Theoretically, our approach departs from conventional hypothesis-testing paradigms and achieves, for the first time, direct adaptation to the degree of misspecification. Empirically, the method substantially improves estimation accuracy and robustness across multiple canonical studies, offering both strong theoretical guarantees and practical computational efficiency.
This paper investigates the sample complexity of Classification Accuracy Testing (CAT) for nonparametric hypothesis testing. We formulate and systematically analyze CAT across three canonical nonparametric settings: discrete distributions, $d$-dimensional distributions with Hölder-continuous densities, and the Gaussian sequence model. CAT trains a binary classifier on synthetic samples and constructs the test statistic from its empirical classification accuracy. We establish, for the first time, that CAT achieves (near) minimax-optimal sample complexity under total variation distance $varepsilon$-separation and type-II error probability $delta$, thereby closing the high-probability complexity gap for likelihood-free testing. Notably, in the discrete two-sample setting, CAT exactly recovers the known minimax-optimal rate. These results provide a theoretically complete foundation for density-ratio-free nonparametric testing.
This paper addresses the optimal approximation of functions in Besov classes under Gaussian noise, measured in the $L_q$-norm. Existing minimax rates fail to precisely capture the dependence on the noise level $sigma$, thereby precluding recovery of the classical noiseless optimal rate as $sigma o 0$. To bridge this theoretical gap, we establish the first noise-level-aware (NLA) sharp minimax convergence rate—providing matching upper and lower bounds and yielding an exact asymptotic characterization of the estimation error in terms of both sample size $m$ and noise variance $sigma^2$. Methodologically, our analysis integrates Besov space theory, information-theoretic lower bound techniques, adaptive truncation, and linear estimation. The resulting rate unifies optimal recovery and minimax estimation frameworks, ensuring a continuous transition to the classical noiseless optimal rate as $sigma o 0$. This provides a unified error benchmark bridging statistical learning and numerical analysis.
This work addresses the limitations of traditional generalization analyses, which rely on the often unverifiable assumption of independent and identically distributed (i.i.d.) data and thus struggle to accurately characterize model performance on unseen data. The paper proposes a deterministic generalization analysis framework that dispenses with any prior probabilistic assumptions. By examining the sensitivity of optimization solutions to data perturbations, it decomposes the generalization error into geometric and probabilistic components, achieving their first-ever decoupling. The framework expresses generalization bounds via a variational principle, leveraging deterministic perturbation analysis and optimization sensitivity theory to capture the discrepancy between in-sample and out-of-sample performance. Error terms are evaluated through posterior statistical hypotheses, enabling the recovery of conventional high-probability or expected generalization guarantees—all without requiring distributional assumptions.
This study addresses the minimax estimation problem under a bounded normal mean model. The authors propose a novel approach based on the stochastic mirror ascent algorithm to approximate the least favorable prior distribution by solving a concave maximization problem, and then adopt its Bayes estimator as an approximate minimax estimator. This work is the first to introduce stochastic mirror ascent into this class of estimation problems and provides theoretical guarantees for the approximation accuracy. Numerical experiments demonstrate that the proposed estimator reduces risk by 6% to nearly 18% compared to the classical minimax linear estimator. Furthermore, the method is successfully applied to impulse response coefficient estimation, confirming its practical effectiveness.
This study addresses the joint problem of composite hypothesis testing and random parameter estimation under distributional uncertainty by establishing a unified minimax optimization framework. Optimal strategies are derived under both Bayesian and Neyman–Pearson–type statistical criteria. The key innovation lies in uncovering an intrinsic connection between the optimal detection–estimation architecture and f-similarity, enabling the identification of least favorable distributions through maximization of f-similarity. Robust design is achieved by integrating this insight with a band-shaped uncertainty model. The proposed approach enhances existing numerical algorithms, offering improved stability while preserving convergence guarantees. Experimental results demonstrate the superiority of the framework in terms of both performance and robustness.
This work addresses the minimax estimation of discrete probability distributions under the $\ell_\infty$ norm, aiming to characterize optimal risk bounds both in expectation and with high probability. By integrating minimax theory, high-dimensional probability analysis, and empirical process techniques with constructive proofs, the study establishes the first fully computable, data-dependent tight risk bound, thereby resolving an open problem posed by Kontorovich and Painsky. The analysis also precisely identifies the structure of extremal distributions that achieve worst-case risk. These theoretical advances not only sharpen existing $\ell_\infty$ risk bounds but also lead to estimators that demonstrate superior empirical performance compared to current methods.
This study establishes minimax lower bounds for the estimation of Maximum Mean Discrepancy (MMD), Hilbert–Schmidt Independence Criterion (HSIC), and Kernelized Stein Discrepancy (KSD) in general topological spaces under unbounded kernel conditions. By integrating reproducing kernel Hilbert space theory, functional analysis, and a minimax information-theoretic framework, the work rigorously proves—under mild assumptions—that the optimal convergence rate for these three classes of kernel-based discrepancy measures remains $n^{-1/2}$. This result resolves a long-standing open theoretical question and extends to the estimation of mean embeddings and centered cross-covariance operators, thereby establishing the minimax optimality of their parametric convergence rates.