finite-sample performance analysis

Designs and analyzes estimators, statistical procedures, and algorithmic dynamics with guarantees that hold for a fixed, finite sample size rather than only asymptotically. Produces non‑asymptotic error bounds, finite‑sample confidence intervals and hypothesis tests, and explicit convergence or performance rates that depend on sample size, dimensionality, and model parameters, including techniques for small‑sample and high‑dimensional regimes.

finite-sampleperformanceanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.39
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work proposes a novel approach to constructing non-asymptotic confidence intervals in randomized experiments that matches the effective sample size of asymptotic methods based on the central limit theorem—a longstanding challenge in finite-sample inference. By systematically leveraging negative dependence and variance-adaptive techniques, the authors derive the first non-asymptotic confidence intervals that achieve the same effective sample size as their asymptotic counterparts. The resulting intervals not only exhibit comparable empirical performance but also attain the information-theoretic lower bound, thereby establishing their optimality for statistical inference with finite samples. This advancement bridges the gap between asymptotic efficiency and finite-sample guarantees, offering a theoretically optimal and practically viable solution for rigorous uncertainty quantification in experimental settings.

effective sample sizenonasymptotic confidence intervalspropensity score

Adaptive A/B Tests and Simultaneous Treatment Parameter Optimization

Oct 13, 2022
YW
Yuhang Wu
🏛️ University of California, Berkeley | Amazon.com Inc

Classical algorithms for strongly convex stochastic optimization achieve fast convergence (O(1/√n)) but suffer from asymptotically non-negligible bias, violating the conditions required for a valid central limit theorem (CLT) and thus impeding asymptotically efficient statistical inference. Method: We propose the first dual-objective algorithm that simultaneously guarantees fast convergence and a provable CLT. Our approach integrates stochastic approximation, asymptotic statistical inference, and adaptive experimental design into a unified framework that ensures asymptotic normality of the estimator. Contribution/Results: We establish theoretical guarantees that the algorithm retains the O(1/√n) convergence rate while satisfying the CLT. Numerical experiments demonstrate substantial improvements over existing methods in estimation accuracy, confidence interval coverage, and identification of optimal treatment parameters. The method provides a new paradigm for continuous, parameterized A/B testing in online platforms—balancing optimization efficiency with statistical reliability.

Addresses non-vanishing bias in stochastic optimization statistical inferenceDevelops algorithm maintaining fast convergence with valid central limit theoremEnables reliable confidence intervals for optimal objective value estimation

In high-dimensional linear regression, existing methods struggle to simultaneously address model selection uncertainty and ensure reliable inference under finite-sample settings. This paper proposes a reproducible-sample-based simulation inference framework that, for the first time, unifies inference for model selection, individual or multiple regression coefficients, and joint parameters—while rigorously guaranteeing finite-sample confidence coverage probability. The method constructs confidence sets via reproducible-sample generation and simulation-based inference, achieving both finite-sample validity and asymptotic optimality. Theoretically, it attains superior coverage accuracy and interval tightness compared to state-of-the-art debiased estimators and bootstrap methods. Empirical evaluations across diverse high-dimensional scenarios confirm its more accurate coverage rates and tighter confidence sets. This work bridges two critical theoretical gaps: (i) valid inference under model selection uncertainty and (ii) finite-sample guarantees in high-dimensional settings.

Addresses model selection uncertainty and guarantees finite-sample performanceConstructs confidence sets for true models and regression coefficientsDevelops finite- and large-sample inference for high-dimensional linear regression

Revisiting Step-Size Assumptions in Stochastic Approximation

May 28, 2024
CK
Caio Kalil Lauand
🏛️ University of Florida

This paper challenges the classical square-summability condition (∑αₙ² < ∞) on step sizes in stochastic approximation (SA), proving for the first time that it is not necessary for convergence. Focusing on power-law step sizes αₙ = α₀n⁻ᵽ with ρ ∈ (0,1) and general Markovian noise, it establishes fine-grained characterizations of convergence, bias, and variance. Key contributions are: (1) the first necessary and sufficient condition for vanishing bias when ρ ≤ 1/2; (2) a proof that Polyak–Ruppert averaging achieves the optimal CLT covariance even for ρ ∈ (0,1/2], though bias dominance slows convergence; and (3) almost-sure and Lₚ convergence for all ρ ∈ (0,1), with an explicit characterization of mean-square error degradation to O(αₙ²). Collectively, these results fundamentally reshape the theoretical foundations of SA step-size design.

Analyzes bias and convergence rates in Markovian noise settings.Challenges traditional step-size assumptions in stochastic approximation.Explores convergence without square-summable step-size conditions.

Latest Papers

What's happening recently
View more

This study addresses the limitations of conventional inference methods rooted in sampling variability when sample sizes approach the population size. By constructing finite populations with known parameters and leveraging CPU/GPU-accelerated repeated sampling experiments, the authors examine the evolution of the randomization distribution of the sample mean across varying sampling fractions. Integrating finite population theory with numerical precision analysis, they demonstrate that in high-coverage scenarios, estimation error predominantly stems from computational precision and architectural constraints rather than sampling randomness. The findings reveal that sampling variability becomes negligible well before exhaustive enumeration is reached, thereby challenging a foundational assumption of classical inferential statistics and offering a basis for rethinking statistical paradigms in the context of large-scale, near-complete data.

finite population samplinglarge-scale datasampling fraction

Design-based finite-sample analysis for regression adjustment

Nov 19, 2025
DS
Dogyoon Song
🏛️ University of California, Davis

This paper addresses finite-sample inference for regression-adjusted average treatment effect (ATE) estimation in high-dimensional randomized experiments where $p > n$. We propose a design-based non-asymptotic analytical framework—distinct from existing approaches relying on correct model specification or large-sample approximations. For the first time, we establish a non-asymptotic theory for high-dimensional regression-adjusted estimators, leveraging Stein’s exchangeable pair technique, Doob martingales, and Freedman’s inequality to characterize how the geometric structure of covariates governs estimation precision. The resulting confidence intervals are instance-adaptive and data-driven, with explicitly computable widths that remain informative even when $p > n$. This work provides the first rigorous, practical, design-based theoretical foundation for robust causal inference in small-sample, high-dimensional settings.

Analyzing regression adjustment for ATE estimation in high-dimensional randomized experimentsDeveloping finite-sample confidence intervals when covariates exceed observationsInvestigating covariate geometry's impact on concentration and bias properties

This work addresses a fundamental challenge in discovery sampling with presence–absence data: determining whether all categories exceeding a given prevalence threshold have been observed. The authors establish the first non-asymptotic, distribution-free, and data-dependent upper confidence bound on the maximum unseen probability—the highest prevalence among unobserved categories—applicable to both bounded and unbounded category spaces, thereby overcoming the limitations of data-independent approaches. Leveraging this bound, they design a sequential stopping rule with finite-sample guarantees and prove its near-optimality via matching upper and lower bounds. The method is developed under a Bernoulli product model through nonparametric inference and worst-case analysis, exhibits robustness to contaminated data, and demonstrates reliable performance in guiding sampling termination decisions across both simulated and real-world datasets.

confidence intervalsdiscovery problemsincidence model

This work proposes a novel asymptotic framework for post-hoc valid inference that overcomes the rigidity of traditional statistical testing, which requires pre-specified significance levels. By extending e-value methodology from non-asymptotic to asymptotic settings, the paper enables the construction of confidence sets and p-values at any significance level chosen after observing the data. The approach achieves sharper and more efficient inference than existing non-asymptotic methods under weaker assumptions—specifically, without requiring strong moment conditions—thereby circumventing the conservativeness inherent in conventional procedures. This advancement establishes a rigorous theoretical foundation for flexible, data-driven statistical analysis while preserving inferential validity in large-sample regimes.

asymptotice-valueslarge-sample

This work addresses the challenge in fixed-confidence best-arm identification, where existing methods often sacrifice either rigorous error control or sample efficiency due to reliance on loose tail bounds or strong parametric assumptions. The authors propose an asymptotic error control framework that constructs anytime-valid confidence sequences tailored for long-horizon experiments, enabling a nonparametric best-arm identification algorithm that leverages individual contextual information. By integrating nonparametric statistical inference with covariate-assisted variance reduction, the method achieves worst-case sample complexity comparable to that of optimal algorithms under Gaussian assumptions with known variance, all under mild regularity conditions. Empirical evaluations demonstrate a substantial reduction in average sample consumption while maintaining strict control over the error rate.

best arm identificationerror controlfixed-confidence

Hot Scholars

OS

Osvaldo Simeone

King's College London
Information theorymachine learningquantum information processingwireless systems
MK

Masahiro Kato

Mizuho-DL Financial Technology Co., Ltd. / The University of Tokyo
Economics
SZ

Shangtong Zhang

University of Virginia
reinforcement learningstochastic approximation
CZ

Changliang Zou

Professor of Statistics, Nankai University
StatisticsQuality Control