high-dimensional asymptotic analysis

Designs and carries out asymptotic and probabilistic analyses that characterize the behavior of estimators, algorithms, spectral objects, or geometric quantities as the number of features, parameters, or dimensions grows large. Builds and applies concentration inequalities, limit theorems, and rate bounds (including large-p and infinite-dimension limits) to quantify identifiability, convergence, phase transitions, and estimation error in high-dimensional regimes.

high-dimensionalasymptoticanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.01
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Spectral Thresholds for Identifiability and Stability:Finite-Sample Phase Transitions in High-Dimensional Learning

Oct 04, 2025
WH
William Hao-Cheng Huang
🏛️ Taiwan Semiconductor Manufacturing Company (TSMC)

In high-dimensional learning, model stability undergoes a sharp phase transition when the sample size $n$ falls below a critical threshold—caused by the weakest Fisher direction being overwhelmed by sampling noise, rendering parameters unidentifiable. To address this, we develop the first non-asymptotic, necessary, and verifiable stability criterion, grounded in the minimal Fisher eigenvalue, and establish a finite-sample phase transition theory that precisely characterizes the phase boundary at the $d/n$ scale. We further propose Fisher-floor regularization—a novel, smoothness- and preprocessing-invariant spectral robustness diagnostic. Integrating Fisher information spectrum analysis, finite-sample random matrix theory, and high-dimensional statistical inference, we empirically validate our framework on Gaussian mixture and logistic regression models: the predicted phase-transition threshold cleanly separates reliable estimation from instability collapse, markedly enhancing interpretability and reliability in high-dimensional modeling.

Bridges theoretical thresholds with practical diagnostics for robust inferenceCharacterizes finite-sample phase transitions in high-dimensional learning stabilityEstablishes spectral threshold for model identifiability using Fisher eigenvalues

Quantitative Error Bounds for Scaling Limits of Stochastic Iterative Algorithms

Jan 21, 2025
XW
Xiaoyu Wang
🏛️ Boston University | ESSEC Business School | University of Waterloo

This work investigates the non-asymptotic pathwise approximation accuracy of stochastic iterative algorithms—such as SGD and SGLD—to the Ornstein–Uhlenbeck process in the univariate setting. Addressing the lack of quantifiable, path-level error bounds in existing theory, we introduce a novel analytical framework for path space by integrating infinite-dimensional Stein’s method with exchangeable pair techniques. This yields explicit convergence rates under both the Lévy–Prokhorov metric and the bounded Wasserstein distance, delivering tight, non-asymptotic upper bounds on pathwise approximation error. We rigorously establish weak convergence and provide quantitative error control for both the iterates’ mean and variance. The framework thus furnishes a foundational toolset for extending the analysis to multivariate settings and more complex stochastic optimization algorithms.

Accuracy and UncertaintyLarge and Complex ProblemsRandomized Iterative Algorithms

Robust Functional Data Analysis for Stochastic Evolution Equations in Infinite Dimensions

Jan 29, 2024
DS
Dennis Schroers
🏛️ Institute of Finance and Statistics | Hausdorff Center for Mathematics | University of Bonn

This paper addresses infinite-dimensional stochastic evolution equations by developing a jump-robust asymptotic theory for covariation differences, aiming to establish a scaling limit relationship between the realized covariation of the solution process and the quadratic covariation of the underlying Hilbert space-valued semimartingale driving process. Methodologically, it integrates infinite-dimensional semimartingale theory, jump-robust estimation, and functional principal component analysis to achieve consistent estimation of the latent driver’s quadratic covariation. Crucially, it introduces and rigorously proves, for the first time, a scaling limit theorem for the covariation difference of the solution process—without requiring continuity of sample paths. The resulting framework enables dynamic-consistent, outlier-robust functional data dimension reduction and stochastic volatility modeling, significantly enhancing robustness and statistical efficiency in high-dimensional functional data settings contaminated by jumps and heavy-tailed noise.

Outlier-robust dimension reduction and volatility model estimationRobust covariation measurement for infinite-dimensional stochastic evolution equationsScaling limits for realized covariations of solution processes

Spectral Estimators for Structured Generalized Linear Models via Approximate Message Passing

Aug 28, 2023
YZ
Yihan Zhang
🏛️ Institute of Science and Technology Austria | University of Cambridge

Parameter estimation in high-dimensional structured generalized linear models suffers from low efficiency, particularly under realistic design matrices exhibiting anisotropy and strong correlations. Method: This paper introduces a novel spectral estimation framework based on Approximate Message Passing (AMP). Contribution/Results: We provide the first exact asymptotic characterization of spectral estimators under correlated Gaussian designs. We identify a universally optimal covariance-adaptive preprocessing strategy, partially resolving a long-standing conjecture on optimal spectral estimation for rotationally invariant models. Theoretically and empirically, our approach substantially reduces sample complexity and achieves provably statistically optimal estimation accuracy—outperforming existing heuristic methods on canonical designs from computational imaging and genomics.

Characterizing spectral estimators for correlated Gaussian designsEstimating parameters in high-dimensional generalized linear modelsIdentifying optimal preprocessing for efficient parameter estimation

Dimension free ridge regression

Oct 16, 2022
CC
Chen Cheng
🏛️ Stanford University | Institute for Advanced Studies | Princeton

This paper investigates the non-asymptotic statistical behavior of ridge regression in high-dimensional and even infinite-dimensional Hilbert spaces, moving beyond the classical proportional-scaling regime of random matrix theory. Methodologically, it integrates tools from random matrix theory, convex concentration inequalities, and spectral analysis to construct an equivalent diagonal sequence model. This enables, for the first time under non-proportional, non-asymptotic scaling, a multiplicative (1±Δ)-approximation of both bias and variance. Key contributions include: (1) dimension-free, explicit non-asymptotic upper bounds on the estimation risk; (2) exact risk characterization for spectrally regular covariates; (3) tight guarantees for benign overfitting in the overparameterized interpolation regime, achieving sharp control of generalization error under zero bias; and (4) unification and refinement of existing proportional asymptotic results.

Characterizes ridge regression for high-dimensional Hilbert covariatesEstablishes non-asymptotic bounds for bias and varianceExtends ridge regression analysis beyond proportional asymptotics

Latest Papers

What's happening recently
View more

This work addresses the limitations of traditional generalization analyses, which rely on the often unverifiable assumption of independent and identically distributed (i.i.d.) data and thus struggle to accurately characterize model performance on unseen data. The paper proposes a deterministic generalization analysis framework that dispenses with any prior probabilistic assumptions. By examining the sensitivity of optimization solutions to data perturbations, it decomposes the generalization error into geometric and probabilistic components, achieving their first-ever decoupling. The framework expresses generalization bounds via a variational principle, leveraging deterministic perturbation analysis and optimization sensitivity theory to capture the discrepancy between in-sample and out-of-sample performance. Error terms are evaluated through posterior statistical hypotheses, enabling the recovery of conventional high-probability or expected generalization guarantees—all without requiring distributional assumptions.

generalizationi.i.d.optimization

This work addresses the challenge in sequential hypothesis testing where model misspecification or estimation error prevents exact construction of e-variables, thereby lacking finite-sample guarantees. We introduce, for the first time, the notion of an asymptotic e-process, defined as a doubly indexed stochastic process $(E_{m,n})$, whose limiting behavior as $m \to \infty$ approximates a standard e-process. We establish its connection to asymptotic supermartingales, derive a corresponding variant of Ville’s inequality, and provide practical construction methods. This framework unifies the theoretical foundation for approximate e-variables, offering sequential inference guarantees under controllable approximation error and explicitly quantifying the trade-off between approximation accuracy and the effective monitoring horizon $r_m$.

asymptotic e-processe-variablesestimation errors

This work addresses the challenge that real-world data in generalized linear models often violate the independent and identically distributed (i.i.d.) assumption. Under the relaxed assumption that the design matrix is orthogonally invariant—meaning its singular vectors are uniformly distributed while singular values remain arbitrary—the paper proposes an efficient parameter estimation method combining optimal spectral initialization with Approximate Message Passing (AMP). The proposed approach achieves the information-theoretically optimal sample complexity for weak recovery and attains the fundamental lower bound on estimation error, thereby extending beyond the classical i.i.d. Gaussian design setting. Rigorous theoretical analysis provides strong performance guarantees, and numerical experiments confirm both the algorithm’s effectiveness and the accuracy of the theoretical predictions on orthogonally invariant as well as more general correlated data.

generalized linear modelsorthogonally invariantparameter estimation

Hot Scholars

LF

Long Feng

Professor of Nankai University
High Dimensional DataHigh Frequency Data
FK

Florent Krzakala

École polytechnique fédérale de Lausanne
Statistical MechanicsStatisticsMachine LearningInformation theory
BL

Bruno Loureiro

École Normale Supérieure & CNRS
Machine LearningStatistical MechanicsDisordered Systems
PZ

Ping Zhao

Hefei University of Technology
Mechanism and RoboticsRehabilitation RoboticsMotion SynthesisComputational Kinematics
LZ

Lenka Zdeborová

EPFL, Switzerland
statistical physicslearning theoryphase transitionsdeep learning