sample complexity analysis

Designs and analyzes statistical sample-size requirements and finite-sample guarantees for learning, estimation, and decision procedures, producing upper and lower sample-complexity bounds, minimax rates, regret bounds, concentration inequalities, and other finite-sample characterizations. Uses probabilistic analysis, probabilistic counting, and output-modeling arguments to derive tight sample-complexity estimates, matching lower bounds and impossibility (e.g., distribution-free or exponential-sample) results, and to quantify how limited data induces error or subgroup disparities.

samplecomplexityanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Improved Compression Bounds for Scenario Decision Making

Jan 15, 2025
GO
Guillaume O. Berger
🏛️ UCLouvain

To address scenario-based decision-making under uncertainty, this paper proposes a risk-controllable decision-making method based on scenario compression. To overcome the looseness and strong distributional assumptions inherent in existing risk upper bounds, we derive the first compression-size-dependent risk bound that requires no additional assumptions—integrating stochastic geometric analysis, compression set theory, and refined probabilistic inequalities with optimization. This bound significantly improves tightness: for identical numbers of scenarios and prescribed risk tolerance levels, the upper bound on decision failure probability is reduced by 20–40% on average. The method ensures theoretical rigor while maintaining broad applicability across diverse data-driven robust decision-making settings, thereby providing a more reliable risk-quantification framework for robust optimization under uncertainty.

Compression TechniquesDecision Making under UncertaintyRisk Mitigation

This work addresses the limitations of traditional generalization analyses, which rely on the often unverifiable assumption of independent and identically distributed (i.i.d.) data and thus struggle to accurately characterize model performance on unseen data. The paper proposes a deterministic generalization analysis framework that dispenses with any prior probabilistic assumptions. By examining the sensitivity of optimization solutions to data perturbations, it decomposes the generalization error into geometric and probabilistic components, achieving their first-ever decoupling. The framework expresses generalization bounds via a variational principle, leveraging deterministic perturbation analysis and optimization sensitivity theory to capture the discrepancy between in-sample and out-of-sample performance. Error terms are evaluated through posterior statistical hypotheses, enabling the recovery of conventional high-probability or expected generalization guarantees—all without requiring distributional assumptions.

generalizationi.i.d.optimization

This work investigates the finite-sample error and computational complexity of Sequential Monte Carlo (SMC) samplers in estimating expectations under a target distribution and its normalizing constant. For both general and annealed sequences of distributions, the study establishes, for the first time, explicit error bounds for standard and waste-free SMC samplers, precisely characterizing their dependence on the number of time steps \(T\) and the ambient dimension \(d\). Leveraging tools from probability theory and statistical learning theory, the authors derive upper bounds that elucidate how estimation error scales with \(T\) and \(d\). Building on these theoretical results, they formulate practical guidelines for algorithm configuration, offering both rigorous theoretical support and actionable recommendations for the efficient deployment of SMC methods in real-world applications.

computational complexityfinite sample boundsnormalising constants

Online Complexity Estimation for Repetitive Scenario Design

Sep 02, 2025
GO
Guillaume O. Berger
🏛️ UCLouvain

To address the challenge of dynamically adapting sample size to risk level (i.e., constraint violation probability) in repeated scenario-based optimization, this paper proposes an online learning method for optimal sample size selection. The method leverages historical scenario solutions and empirically observed violation probabilities to estimate the risk distribution function in real time, thereby establishing a nonlinear mapping between sample size and risk. It is the first approach to achieve online, adaptive estimation of the optimal sample size under non-fixed computational complexity, nonconvex constraints, and time-varying distributions—overcoming the reliance of conventional scenario optimization on static problem structures—and provides theoretical guarantees of convergence. Experiments demonstrate significant improvements in both risk control accuracy and computational efficiency across diverse challenging scenarios, validating its applicability to repetitive decision-making tasks such as power system dispatch and financial risk management.

Ensuring convergence for fixed-complexity scenario optimization problemsEstimating optimal sample size for repetitive scenario designLearning risk probability density function from observed data

Fast Rate Information-theoretic Bounds on Generalization Errors

Mar 26, 2023
XW
Xuetong Wu
🏛️ University of Melbourne

This work addresses the tightness of information-theoretic generalization error bounds with respect to sample size $n$, particularly the looseness of the individual-sample mutual information (ISMI) bound. To overcome the suboptimal $O(1/sqrt{n})$ convergence rate, we introduce, for the first time, an *excess risk assumption*, yielding a tight $O(1/n)$ fast-rate bound. Furthermore, we propose a novel generalization framework based on the $(eta,c)$-central condition, under which the mutual information term directly governs the convergence rate. We rigorously prove that this bound achieves the optimal $O(1/n)$ rate under standard assumptions. Empirical evaluation on canonical tasks—such as Gaussian mean estimation—demonstrates substantial improvements over existing information-theoretic bounds. The proposed framework thus bridges theoretical rigor with practical superiority, advancing both the tightness and applicability of information-theoretic generalization analysis.

Investigates tightness of generalization error boundsProposes new bounds using (η, c)-central conditionShows fast rate recovery under excess risk assumption

Latest Papers

What's happening recently
View more

This work addresses the lack of non-asymptotic sample complexity guarantees for learning high-dimensional continuous exponential family distributions, particularly in the challenging setting of unbounded support. By employing score matching to perform structure learning on polynomial-form exponential family models, the study establishes the first finite-sample error bounds for this class of models through a synthesis of high-dimensional probabilistic modeling and non-asymptotic statistical analysis. The results demonstrate that the required sample size scales polynomially with the ambient dimension, thereby providing the first rigorous sample complexity guarantee for learning high-dimensional exponential family distributions with unbounded support. This contribution fills a critical theoretical gap in the non-asymptotic understanding of continuous exponential families.

exponential familieshigh-dimensional statisticsnon-asymptotic bounds

This work investigates the fundamental performance limits of learning and estimation tasks within an information-theoretic framework, independent of the computational capabilities of specific algorithms. By integrating tools from information theory and statistical learning theory—including metric entropy, VC dimension, Rademacher complexity, mutual information, and relative entropy—it systematically derives multiple upper bounds on generalization error. Simultaneously, leveraging Fano’s inequality together with covering and packing numbers, the study establishes information-theoretic lower bounds on minimax risk. The analysis unifies two complementary paradigms: one grounded in the geometric structure of metric spaces and the other based on information-theoretic measures. This synthesis yields a rigorous and broadly applicable theoretical framework for characterizing the optimal performance boundaries inherent to learning and estimation problems.

estimationgeneralization errorinformation-theoretic limits

This work addresses the challenge of generalization learning under highly dependent, non-i.i.d. data by introducing a learning framework based on simulatable processes, wherein the learner has access to a simulator that approximates the true data-generating mechanism. The authors propose a unified algorithm—enabled by the novel incorporation of time-bounded Kolmogorov complexity—that simultaneously learns any hypothesis class with finite VC dimension. They rigorously establish the statistical and computational advantages of conditional sampling within this setting. The approach achieves no-regret learning across all polynomial-time simulatable processes, with error bounds depending solely on the VC dimension, thereby substantially extending the applicability of the classical PAC learning model.

dependent datageneralizationlearning theory

This study addresses the joint optimization of sampling design and estimation under bounded population values, aiming to achieve design-unbiased estimation of the population total while minimizing the worst-case mean squared error. Within the Horvitz–Thompson framework and adopting a minimax criterion, the paper establishes—for the first time—the minimax lower bound over all design-unbiased estimators and shows that this bound is attainable when the unit inclusion indicators are pairwise independent. The authors further propose a midpoint-differenced Horvitz–Thompson estimator, which achieves minimax optimality under an independent sampling strategy with inclusion probabilities πᵢ* = min(1, c(bᵢ − aᵢ)). This estimator is also shown to be admissible within the class of unbiased and affine-equivariant estimators, thereby extending Gabler’s (1990) linear result to a broader class of estimators.

bounded outcomesdesign-unbiased estimationfinite population

Hot Scholars

ID

Ilias Diakonikolas

University of Wisconsin-Madison
theoretical computer sciencealgorithmic statisticsmachine learningprobability theory
SM

Shay Moran

(Math, CS, & DDS, Technion) & (Google Research)
Computer ScienceMathematics
EM

Eric Moulines

Professeur, Ecole Polytechnique, Membre de l'Académie des Sciences
StatisticsMachine learningSignal Processing
SS

Sergey Samsonov

HSE university, Moscow
high-dimensional probabilityMarkov ChainsMCMC
AN

Alexey Naumov

Professor, HSE University
probability theorystatisticsmachine learningrandom matrices