Score
Derives provable information-theoretic and probabilistic bounds that quantify performance limits of estimators, tests, or learning algorithms — for example, using Fano’s inequality and mutual information to obtain lower bounds on error or sample complexity, and using concentration/tail inequalities, union-bound techniques, total-variation or spectral-norm arguments to obtain probabilistic error, regret, or minimax rate bounds. These derivations produce concrete dimension-dependent rates, sample-complexity or regret statements, and tail-probability guarantees for concrete models or procedures.
This work addresses the tightness of information-theoretic generalization error bounds with respect to sample size $n$, particularly the looseness of the individual-sample mutual information (ISMI) bound. To overcome the suboptimal $O(1/sqrt{n})$ convergence rate, we introduce, for the first time, an *excess risk assumption*, yielding a tight $O(1/n)$ fast-rate bound. Furthermore, we propose a novel generalization framework based on the $(eta,c)$-central condition, under which the mutual information term directly governs the convergence rate. We rigorously prove that this bound achieves the optimal $O(1/n)$ rate under standard assumptions. Empirical evaluation on canonical tasks—such as Gaussian mean estimation—demonstrates substantial improvements over existing information-theoretic bounds. The proposed framework thus bridges theoretical rigor with practical superiority, advancing both the tightness and applicability of information-theoretic generalization analysis.
This work addresses the lack of information-theoretic lower bounds for tail-sensitive objectives—such as Conditional Value-at-Risk (CVaR)—in interactive statistical decision-making, where existing bounds primarily target expected risk. The authors propose a generalized Fano framework that replaces the deterministic success event in classical Fano inequalities with a randomized one-bit statistic derived from an arbitrary bounded transformation of the loss function. By integrating Bernoulli f-divergences, the Rockafellar–Uryasev representation of CVaR, and a Pinsker-type inequality, they establish the first information-theoretic lower bound on Bayesian CVaR under bounded losses. This bound is explicitly expressed in terms of mutual information and extends prior interactive Fano results under both KL divergence and mixture reference distributions, thereby providing a theoretical foundation for tail-risk optimization.
This work addresses the precise characterization of information-theoretic limits and convergence rates in Bayesian nonparametrics. Focusing on fractional posteriors, variational inference, and maximum likelihood estimation, we derive a universal upper bound on mutual information that uniformly governs their statistical convergence rates. Our bound substantially improves existing contraction rate guarantees for fractional posteriors and, for the first time, establishes a rigorous theoretical bridge between the PAC-Bayes framework and Bayesian nonparametric information theory. The results yield tighter, quantifiable convergence rates and unify disparate inferential paradigms under a single information-theoretic analytical framework. This provides a principled, comparable benchmark for fundamental statistical limits in nonparametric settings.
This work addresses the problem of deriving generalization error upper bounds for batch learning algorithms under mixing stochastic processes (i.e., dependent data), without imposing any stability assumptions on the batch learner. The method introduces a novel analytical framework based on Online-to-Batch conversion: stability requirements are shifted to an associated online learner, and a new notion of online algorithm stability—defined via the first-order Wasserstein distance—is proposed for the first time. It is shown that the Exponentially Weighted Average (EWA) algorithm satisfies this stability condition. Consequently, the framework yields both expectation- and high-probability generalization bounds applicable to *any* batch learning algorithm; under i.i.d. assumptions, the bounds reduce to classical forms up to correction terms governed by the mixing decay rate. The resulting bounds are explicitly computable, substantially broadening the applicability and practical utility of learning theory under data dependence.
High-probability analysis of learning algorithms involving light-tailed (e.g., sub-exponential, sub-Gaussian) but possibly unbounded random variables poses significant technical challenges due to the lack of uniform concentration tools across distribution families. Method: We propose a generic black-box reduction that systematically transforms high-probability analysis of any algorithm relying on light-tailed randomness into the corresponding analysis under bounded-variable assumptions, incurring only controllable logarithmic-factor overheads. Contribution/Results: This is the first unified framework handling diverse light-tailed distributions without ad hoc concentration inequalities—greatly simplifying theoretical analysis. As applications, we reconstruct a generalized Azuma’s inequality and derive tight high-probability convergence bounds for stochastic optimization algorithms under light-tailed noise, demonstrating both the method’s effectiveness and broad applicability.
Classical learning theory, relying on uniform convergence over hypothesis spaces, struggles to explain the strong generalization performance of over-parameterized deep neural networks. This work proposes a unified framework that, for the first time, incorporates data-dependent worst-case generalization bounds into a single template inequality. The framework systematically integrates extensions of PAC-Bayes theory, geometric and topological characterizations of optimization trajectories—such as fractal dimension and α-weighted persistence sums—and information-theoretic surrogates grounded in algorithmic stability. By enabling direct comparison among diverse generalization bounds, it reveals their intrinsic connections and distinctions, and establishes a non-vacuous, tight family of upper bounds on generalization error that effectively accounts for the empirical generalization behavior of over-parameterized models.
This study investigates the fundamental structural limitations governing the predictive performance of machine learning–based decision systems, demonstrating that these constraints arise from intrinsic properties of the data-generating process rather than algorithmic choices. By integrating information-theoretic bounds (Fano’s and Cramér–Rao inequalities), models of interaction dependencies (Markov random fields and potential functions), and feedback-driven stochastic dynamical systems—including those incorporating large language model agents—the work establishes a unified analytical framework. This framework reveals how structural assumptions such as independence and ergodicity critically determine the validity of statistical inference. The analysis proves that the ultimate ceiling on predictive capability is dictated by the underlying data-generation mechanism, thereby providing a theoretical foundation for designing reliable decision systems that respect these inherent structural constraints.
This study investigates high-confidence non-asymptotic upper and lower bounds for the minimal risk in statistical learning, circumventing the conventional reliance on boundedness assumptions of the empirical risk function. By integrating sharp forms of Talagrand’s concentration inequality—specifically the Bousquet and Klein–Rio refinements—with transport-entropy inequalities and empirical process theory, the authors derive a lower bound independent of both the number of parameters and input dimension, under Gaussian or exponential integrability assumptions. The corresponding upper bound is characterized by the interplay between sample size and the box-counting dimension of the parameter set measured in an Orlicz norm. This work thus provides a more general and non-asymptotic theoretical framework for evaluating learning algorithm performance without resorting to asymptotic approximations.
This work addresses the minimax estimation of discrete probability distributions under the $\ell_\infty$ norm, aiming to characterize optimal risk bounds both in expectation and with high probability. By integrating minimax theory, high-dimensional probability analysis, and empirical process techniques with constructive proofs, the study establishes the first fully computable, data-dependent tight risk bound, thereby resolving an open problem posed by Kontorovich and Painsky. The analysis also precisely identifies the structure of extremal distributions that achieve worst-case risk. These theoretical advances not only sharpen existing $\ell_\infty$ risk bounds but also lead to estimators that demonstrate superior empirical performance compared to current methods.