Score
Derive theoretical guarantees that quantify how a model or learning algorithm’s performance on finite training data transfers to expected (population) performance by proving sample complexity and generalization error bounds. This involves constructing and analyzing proofs that produce explicit dependence on the model class, algorithm, training procedure, and data (e.g., via concentration inequalities, uniform convergence, Rademacher/VC complexity, stability, PAC‑Bayes, covering numbers), and reporting the resulting learning bounds and their assumptions.
Classical learning theory, relying on uniform convergence over hypothesis spaces, struggles to explain the strong generalization performance of over-parameterized deep neural networks. This work proposes a unified framework that, for the first time, incorporates data-dependent worst-case generalization bounds into a single template inequality. The framework systematically integrates extensions of PAC-Bayes theory, geometric and topological characterizations of optimization trajectories—such as fractal dimension and α-weighted persistence sums—and information-theoretic surrogates grounded in algorithmic stability. By enabling direct comparison among diverse generalization bounds, it reveals their intrinsic connections and distinctions, and establishes a non-vacuous, tight family of upper bounds on generalization error that effectively accounts for the empirical generalization behavior of over-parameterized models.
This work addresses the tightness of information-theoretic generalization error bounds with respect to sample size $n$, particularly the looseness of the individual-sample mutual information (ISMI) bound. To overcome the suboptimal $O(1/sqrt{n})$ convergence rate, we introduce, for the first time, an *excess risk assumption*, yielding a tight $O(1/n)$ fast-rate bound. Furthermore, we propose a novel generalization framework based on the $(eta,c)$-central condition, under which the mutual information term directly governs the convergence rate. We rigorously prove that this bound achieves the optimal $O(1/n)$ rate under standard assumptions. Empirical evaluation on canonical tasks—such as Gaussian mean estimation—demonstrates substantial improvements over existing information-theoretic bounds. The proposed framework thus bridges theoretical rigor with practical superiority, advancing both the tightness and applicability of information-theoretic generalization analysis.
This work addresses the problem of deriving generalization error upper bounds for batch learning algorithms under mixing stochastic processes (i.e., dependent data), without imposing any stability assumptions on the batch learner. The method introduces a novel analytical framework based on Online-to-Batch conversion: stability requirements are shifted to an associated online learner, and a new notion of online algorithm stability—defined via the first-order Wasserstein distance—is proposed for the first time. It is shown that the Exponentially Weighted Average (EWA) algorithm satisfies this stability condition. Consequently, the framework yields both expectation- and high-probability generalization bounds applicable to *any* batch learning algorithm; under i.i.d. assumptions, the bounds reduce to classical forms up to correction terms governed by the mixing decay rate. The resulting bounds are explicitly computable, substantially broadening the applicability and practical utility of learning theory under data dependence.
This work aims to provide a more precise characterization of the generalization error of differentially private learning algorithms. By integrating information-theoretic tools with typicality analysis and leveraging the stability properties inherent to differential privacy, we rigorously refine existing mutual information-based upper bounds and establish, for the first time, a computable and tight upper bound on maximal leakage. The proposed approach applies uniformly to both expected and high-probability analyses of generalization error. The resulting bounds are not only tighter than prior results but also exhibit strong computability, thereby offering enhanced guarantees on generalization performance.
This work addresses the family of parametric optimization problems and proposes the first unified, data-driven framework for analyzing the generalization performance of both classical and learned optimizers. Methodologically: (1) it introduces PAC-Bayes theory to the analysis of learned optimizers, deriving verifiable, high-probability generalization upper bounds; (2) it establishes performance bounds for classical optimizers based on empirical convergence rates; and (3) it pioneers a learning paradigm that directly minimizes the PAC-Bayes bound during training. Evaluated on signal processing, control, and meta-learning tasks, the derived bounds are significantly tighter than conventional worst-case guarantees. Moreover, the theoretical generalization guarantees for learned optimizers consistently exceed the empirical performance of their non-learned baselines—thereby unifying theoretical rigor with practical efficacy.
In high-dimensional, small-sample regimes, conventional machine learning generalization bounds suffer from excessive looseness due to the “curse of dimensionality.” Method: This paper proposes an adaptive family of generalization bounds tailored to discretized Euclidean spaces. We first derive a non-asymptotic concentration inequality for finite metric spaces; introduce geometric representation dimension (m) as a pivotal parameter; and construct bounds with constant factor (c_m) scaling as (sqrt{m}). Tight analysis is achieved via metric embedding combined with discretization-based modeling. Contribution/Results: The proposed bounds yield significant tightening under practical sample sizes; retain the optimal (O(1/sqrt{N})) convergence rate; and break the exponential or polynomial dependence on ambient dimension inherent in classical bounds—achieving constant-factor improvement even in high-dimensional, low-precision settings.
This work addresses the limitations of traditional generalization analyses, which rely on the often unverifiable assumption of independent and identically distributed (i.i.d.) data and thus struggle to accurately characterize model performance on unseen data. The paper proposes a deterministic generalization analysis framework that dispenses with any prior probabilistic assumptions. By examining the sensitivity of optimization solutions to data perturbations, it decomposes the generalization error into geometric and probabilistic components, achieving their first-ever decoupling. The framework expresses generalization bounds via a variational principle, leveraging deterministic perturbation analysis and optimization sensitivity theory to capture the discrepancy between in-sample and out-of-sample performance. Error terms are evaluated through posterior statistical hypotheses, enabling the recovery of conventional high-probability or expected generalization guarantees—all without requiring distributional assumptions.
This work investigates the generalization ability and sample complexity of supervised learning from an information-theoretic perspective, modeling the learning process as lossy compression with finite blocklength: training data sampling corresponds to encoding, and model construction to decoding. It introduces, for the first time, finite-blocklength lossy compression theory to analyze generalization error, deriving fundamental lower bounds on both generalization error and sample complexity for any fixed randomized learning algorithm and its optimal sampling strategy. The proposed framework cleanly disentangles the distinct contributions of overfitting and task-inductive bias mismatch, while unifying information-theoretic generalization bounds with the algorithmic stability perspective, thereby revealing their essential roles in determining generalization performance.
This work addresses the challenge of generalization learning under highly dependent, non-i.i.d. data by introducing a learning framework based on simulatable processes, wherein the learner has access to a simulator that approximates the true data-generating mechanism. The authors propose a unified algorithm—enabled by the novel incorporation of time-bounded Kolmogorov complexity—that simultaneously learns any hypothesis class with finite VC dimension. They rigorously establish the statistical and computational advantages of conditional sampling within this setting. The approach achieves no-regret learning across all polynomial-time simulatable processes, with error bounds depending solely on the VC dimension, thereby substantially extending the applicability of the classical PAC learning model.
This work addresses the limitations of classical algorithmic stability theory, which typically relies on strong assumptions such as bounded loss functions or sub-Gaussian/sub-Weibull tail behavior—conditions often violated in heavy-tailed or unbounded loss settings. The paper introduces a novel $L_p$ stability framework that requires only finite $L_p$ moments of the loss function, thereby relaxing the conventional bounded differences condition. By extending McDiarmid’s inequality to accommodate $L_p$ constraints, the authors derive sharp high-probability generalization bounds under this significantly weaker assumption. This approach is shown to be broadly applicable across multiple learning paradigms, including empirical risk minimization, transductive regression, and meta-learning, demonstrating that robust generalization guarantees can still be achieved even when losses are unbounded, provided $L_p$ stability holds.
This work investigates the fundamental performance limits of learning and estimation tasks within an information-theoretic framework, independent of the computational capabilities of specific algorithms. By integrating tools from information theory and statistical learning theory—including metric entropy, VC dimension, Rademacher complexity, mutual information, and relative entropy—it systematically derives multiple upper bounds on generalization error. Simultaneously, leveraging Fano’s inequality together with covering and packing numbers, the study establishes information-theoretic lower bounds on minimax risk. The analysis unifies two complementary paradigms: one grounded in the geometric structure of metric spaces and the other based on information-theoretic measures. This synthesis yields a rigorous and broadly applicable theoretical framework for characterizing the optimal performance boundaries inherent to learning and estimation problems.