high-dimensional probability

Deriving probabilistic limit theorems and finite-sample characterizations for multivariate/high-dimensional random variables and statistics (e.g., multivariate CLTs, extreme-value limits, Taylor expansions and state-evolution analyses), including establishing probabilistic thresholds and dimensionality regimes (relations between p and n) under which asymptotic or finite-sample guarantees hold.

high-dimensionalprobability

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Taylor's Theorem and Mean Value Theorem for Random Functions and Random Variables

Feb 20, 2021
YY
Yifan Yang
🏛️ Case Western Reserve University | University of Maryland

Classical Taylor’s theorem—developed in a deterministic setting—lacks measurability guarantees for the intermediate point when applied to random functions (e.g., likelihoods), undermining the rigorous probabilistic interpretation of asymptotic expansions in statistics. This foundational issue has long been overlooked. Method: Under mild measurability assumptions, we establish the first multivariate Taylor and mean-value theorems applicable to random vectors and random functions, explicitly ensuring the measurability of the intermediate point. Our approach integrates stochastic processes, measure theory, multivariate differential calculus, and statistical asymptotics. Contribution/Results: The results bridge a critical theoretical gap between stochastic analysis and statistical inference. They provide a strictly measurable foundation for asymptotic expansions of estimators—including maximum likelihood and M-estimators—thereby strengthening the mathematical rigor of statistical inference and enabling valid probabilistic reasoning in high-order approximations.

Ensuring measurability of intermediate points in random functionsExtending Taylor's theorem to stochastic functions and variablesProviding rigorous foundation for Taylor expansions in statistics

This paper addresses the conditional distribution of Banach space-valued jointly Gaussian random variables. It establishes that the conditional distribution remains Gaussian and develops a finite-dimensional approximation scheme based on Banach space-valued martingales to compute the conditional mean and covariance operator exactly. Methodologically, it unifies nuclear norm convergence and weak convergence analyses—yielding, for the first time, a rigorously convergent Gaussian conditioning theory in general Banach spaces. The framework applies broadly, including to reproducing kernel Hilbert spaces (RKHS) and spaces of continuous functions. For continuous Gaussian process paths, it guarantees uniform convergence of the conditional mean and covariance functions, as well as weak convergence of the conditional probability measures. These results provide a rigorous, general mathematical foundation for infinite-dimensional statistical inference and Bayesian inverse problems involving Gaussian processes.

Application to Gaussian processes in machine learningApproximation scheme for means and covariancesConditional distributions of Gaussian random variables

Generation through the lens of learning theory

Oct 17, 2024
VR
Vinod Raman
🏛️ University of Michigan

This paper investigates the theoretical foundations of generative capability, formalizing necessary and sufficient conditions for decidability of generation—unified, non-uniform, and prompt-based—and elucidating its fundamental relationship with predictability. Building on statistical learning theory, we introduce *closure dimension*—a novel combinatorial measure—as the core criterion characterizing generative capacity, thereby establishing a general framework for generability. We rigorously prove that generability and predictability (in both PAC and online learning models) are mutually exclusive. We provide complete, precise decidability characterizations—necessary and sufficient conditions—for all three generation paradigms. Furthermore, we unify and extend the framework of Kleinberg & Mullainathan (2024), modeling generation as a formal language closure problem over binary hypothesis classes. Our results bridge foundational learning theory with modern generative modeling, offering rigorous criteria to distinguish generative from predictive computation.

Generative ConditionsLearning TheoryPredictivity

Generalized Universal Inference on Risk Minimizers

Jan 31, 2024
ND
N. Dey
🏛️ North Carolina State University

This paper addresses uncertainty quantification for risk-minimizing estimators in machine learning, overcoming limitations of classical approaches that rely on restrictive distributional assumptions and asymptotic theory. We propose the first general-purpose, finite-sample, distribution-free, and frequentist-valid inference framework applicable to *any* risk minimizer. Our method is grounded in the generalized likelihood ratio test, integrated with empirical process analysis and data-driven tuning, and inherently supports anytime-valid inference. Theoretically, it guarantees exact coverage of confidence sets for *all* finite sample sizes—without asymptotic approximations. Empirically, it consistently outperforms classical asymptotic methods across diverse tasks, demonstrating both high accuracy and strong robustness. This work establishes a new paradigm for model-agnostic statistical inference.

Achieve finite-sample validity in statistical learningEstimate unknowns with uncertainty quantificationGeneralize universal inference for risk minimizers

This work addresses posterior functional computation in high-dimensional generalized linear models under non-log-concave likelihoods. To overcome the lack of non-asymptotic theoretical guarantees for conventional MCMC methods under non-convex posteriors, we propose a gradient-driven MCMC algorithm that achieves statistically optimal posterior sampling in polynomial time, requiring only local likelihood regularity and a suitable initial point. Our contribution is the first non-asymptotic convergence theory for posterior sampling that operates outside the M-estimation framework—free from asymptotic assumptions and scalable to high dimensions—applicable to density estimation, nonparametric regression, and PDE inverse problems. Crucially, the theoretical guarantees hold rigorously under non-log-concave likelihoods. Empirical evaluations confirm the method’s effectiveness and computational efficiency in both generalized linear models and PDE inverse problems.

Computing posterior functionals in high-dimensional modelsNon-log-concave likelihood functions in statistical modelsPolynomial time guarantees for gradient based MCMC

Latest Papers

What's happening recently
View more

This work addresses the absence of a unified theoretical framework for the Central Limit Theorem (CLT), which has led to fragmented proofs and obscured insights into its intrinsic structure. To resolve this, the authors propose an abstract framework grounded in enriched category theory with expansive seminorms, uniquely integrating categorical methods with expansion structures to systematically characterize and derive diverse CLTs. This approach not only recovers and strengthens classical results—including the standard CLT and the Law of Large Numbers—but also extends them to statistical mechanical contexts, yielding a novel CLT for observables on symplectic manifolds. The framework thus establishes a categorical foundation for probabilistic limit theorems and opens avenues for broad theoretical and applied generalizations.

abstract probabilityCentral Limit Theoremlimit theorems

This work addresses the efficient construction and utilization of Stein methods for distributional metrics and optimization in probabilistic reasoning and learning. By establishing a general framework for constructing Stein discrepancies, it systematically analyzes their computability, separating properties, and capacity to detect and control convergence of probability measures. The study further elucidates the theoretical connection between Stein operators and Stein Variational Gradient Descent (SVGD). Beyond clarifying the foundational theoretical properties of Stein discrepancies, this research develops a unified framework that bridges variational inference and nonparametric distribution approximation, thereby providing a rigorous theoretical foundation and practical pathway for designing efficient and verifiable probabilistic learning algorithms.

convergence controlprobabilistic inferenceStein discrepancy

This work investigates the minimax detection boundary for uniformity testing of multinomial distributions under ℓₚ deviations, focusing on the intermediate regime where the sample size satisfies \( N = o(n^2) \) and the signal-to-noise ratio converges to a finite positive constant. By introducing a Poissonized model, constructing a Poisson mixture prior, and applying a conditional central limit theorem for weighted sums, the authors establish—for the first time—a matching lower bound on the minimax risk in this regime. This lower bound precisely coincides with the known upper bound, thereby fully characterizing the asymptotic behavior of the minimax risk at the level of sharp constants: the risk converges to the nontrivial constant \( 2\Phi(-u^*/2) \).

goodness-of-fitintermediate regimeminimax risk

This study investigates the large-sample asymptotic behavior of the frequency spectrum in Pitman–Yor random partitions, with particular emphasis on the limiting distribution of the sum of frequency counts over intervals of the form ∑_{j=⌊λn⌋}^{⌊μn⌋} M_{jn}. By leveraging the theory of Gibbs-type partitions, asymptotic analysis, and methods from combinatorial stochastic structures, the work establishes, for the first time, a limit theorem for the Pitman–Yor frequency spectrum that holds uniformly across broad ranges of such intervals, and further explores its functional limit form. The results uncover a profound connection between the frequency spectrum and the limiting shape of associated random combinatorial structures, thereby providing a rigorous theoretical foundation for applications in population genetics, particularly in the analysis of allele frequency spectra.

asymptotic distributionfrequency spectrumGibbs-type partitions

This study investigates the asymptotic normality of pattern counts in random planar maps. By directly analyzing the bivariate coefficient asymptotics of functional equations involving a single catalytic variable, the authors establish a central limit theorem for pattern occurrences without relying on prior assumptions about face-degree distributions. The approach leverages techniques from analytic combinatorics—specifically, functional equations and bivariate asymptotic analysis—to yield a more streamlined proof while extending the result to arbitrary boundary conditions and broader classes of maps. This work substantially broadens the scope under which asymptotic normality holds, enhancing both the generality and technical efficiency of the underlying theory.

asymptotic normalitypattern countsplanar maps

Hot Scholars

AR

Aaditya Ramdas

Associate Professor (with tenure), Carnegie Mellon University
Machine LearningStatistics
LF

Long Feng

Professor of Nankai University
High Dimensional DataHigh Frequency Data
NL

Nils Lid Hjort

Professor of Mathematical Statistics, University of Oslo
Theoretical and applied statistics and probability theory
RW

Ruodu Wang

University of Waterloo
StatisticsRisk ManagementActuarial ScienceFinancial Engineering
HL

Han Lin Shang

Department of Actuarial Studies and Business Analytics, Macquarie University
Functional data analysisnonparametric smoothingnonparametric statisticsmachine learning