Score
Theoretical and applied analysis of spectral properties of large random matrices, including eigenvalue/eigenvector perturbation and asymptotic guarantees under noise. It is used to characterize geometric biases from heteroskedastic noise, derive diagnostics for model failures, and justify estimators that shrink spurious components.
To address the challenge of accurately estimating eigenvalues of precision matrices in high-dimensional settings—where direct inversion of the sample covariance matrix is ill-posed—the paper proposes a novel, inversion-free estimator. Leveraging random matrix theory, it rigorously analyzes the convergence rates of the Stieltjes transform and its derivative of the sample covariance matrix, enabling the construction of a refined eigenvalue estimator. The estimator achieves an asymptotic bias of order $O(1/K^2)$, markedly improving upon the conventional $O(1/K)$ rate; moreover, a rigorous central limit theorem is established, precisely characterizing the asymptotic distribution of the estimator. Theoretical analysis confirms consistency under the high-dimensional asymptotic regime ($n, p o infty$, $p/n o c > 0$), while numerical experiments validate both the enhanced estimation accuracy and the fidelity of the limiting distribution. This work provides the first direct estimation framework for spectral inference of high-dimensional precision matrices that simultaneously achieves higher-order bias correction and verifiable asymptotic normality.
This paper addresses spectral analysis of the instantaneous volatility matrix for high-dimensional continuous-time processes under high-frequency observation, overcoming the classical random matrix theory’s reliance on i.i.d. samples. Methodologically, it extends large-dimensional covariance matrix spectral theory to instantaneous volatility matrices, establishing—for the first time—the first-order limit (a Marchenko–Pastur-type limit) of their empirical spectral distribution and the central limit theorem for linear spectral statistics. Based on these theoretical foundations, it constructs novel identity and sphericity tests tailored for high-dimensional high-frequency data. The analysis integrates large-dimensional asymptotics, high-frequency sampling theory, and the limit theory of linear spectral statistics. Extensive simulations confirm the robustness and finite-sample validity of the proposed test statistics. The work provides a new theoretical framework and practical inferential tools for high-dimensional volatility structure analysis in finance and related fields.
Parameter estimation in high-dimensional structured generalized linear models suffers from low efficiency, particularly under realistic design matrices exhibiting anisotropy and strong correlations. Method: This paper introduces a novel spectral estimation framework based on Approximate Message Passing (AMP). Contribution/Results: We provide the first exact asymptotic characterization of spectral estimators under correlated Gaussian designs. We identify a universally optimal covariance-adaptive preprocessing strategy, partially resolving a long-standing conjecture on optimal spectral estimation for rotationally invariant models. Theoretically and empirically, our approach substantially reduces sample complexity and achieves provably statistically optimal estimation accuracy—outperforming existing heuristic methods on canonical designs from computational imaging and genomics.
This paper investigates the statistical stability of randomized singular value decomposition (RSVD) under a signal-plus-noise model, focusing on ℓ₂ and ℓ_{2,∞} errors between approximate and true left singular vectors, as well as entrywise errors after projection. Methodologically, it integrates Gaussian random sketching, power iteration acceleration, and refined perturbation analysis. The contributions are threefold: (i) it establishes the first error bounds explicitly dependent on the signal-to-noise ratio (SNR); (ii) it characterizes a sharp phase-transition threshold for the number of power iterations (g) governing estimation accuracy; and (iii) it proves row-wise and entrywise asymptotic normality of the RSVD estimators. These results provide near-optimal theoretical guarantees for community detection, PCA with missing data, and matrix completion. Crucially, the derived error bounds quantitatively reveal the synergistic interplay between SNR and iteration count—highlighting how increased iterations mitigate low-SNR degradation but exhibit diminishing returns beyond the phase transition.
This paper investigates the generalization performance of random feature methods under generalized spectral regularization—including explicit schemes (e.g., Tikhonov regularization) and implicit schemes (e.g., gradient descent and accelerated algorithms). Leveraging insights from neural tangent kernel (NTK) theory, it establishes optimal learning rates for the first time under non-RKHS-type regularity assumptions, thereby unifying the analysis of implicit and explicit regularization mechanisms. Methodologically, the work integrates random feature mappings, spectral analysis, source condition modeling, and rigorous generalization error bound derivation to obtain tight convergence rates for both neural networks and neural operators. Key contributions are: (1) extending optimal convergence guarantees beyond the RKHS framework to broader classes of spectral regularity; (2) providing a unified, tight, and improved theoretical bound applicable to diverse kernel-based algorithms; and (3) substantially enhancing computational efficiency and depth of generalization analysis for large-scale kernel methods.
Classical spectral perturbation theory, such as the Davis–Kahan theorem, fails to capture the systematic geometric bias in the leading eigenspace under heterogeneous noise. This work studies a signal-plus-noise model with heteroscedastic and sparse random perturbations, and for the first time identifies and quantifies a deterministic geometric bias induced by the alignment between the signal structure and the noise variance profile. The total perturbation is decomposed into three components: a signal-to-noise ratio term, a stochastic fluctuation term, and this geometric bias term. Leveraging the quadratic vector equation (QVE) framework and refined isotropic local laws, the paper establishes non-asymptotic perturbation bounds in both operator norm and ℓ²→ℓ^∞ norm. The resulting upper bounds are nearly optimal and substantially outperform classical theory in predicting eigenspace perturbations under heterogeneous noise.
This study addresses the frequent misinterpretation of fluctuations in dominant eigensubspaces and scalar spectral functionals—such as the absorption ratio—in rolling covariance estimation as genuine market structural changes, when they are often artifacts of estimation noise, particularly under shrinkage. By leveraging perturbation analysis and calibrated inference, the work derives, for the first time, the first-order null distribution of eigensubspace variation under overlapping windows and establishes its invariance under rotation-equivariant shrinkage estimators. It further shows that only scale-invariant spectral functionals enjoy first-order immunity to elliptical kurtosis. To correct high-dimensional bias in the absorption ratio, a trace-preserving spiked debiased estimator is proposed. Theoretical results, supported by Davis–Kahan bounds, distribution-free confidence bands, and an estimator-aware bootstrap, are validated through simulations and successfully applied to equity data for reliable detection of true market structural shifts.
This work addresses the failure of conventional quadratic forms involving precision matrices in high-dimensional settings where the feature dimension \(p\) exceeds the sample size \(n\), a regime plagued by rank deficiency and excessive complexity. The authors propose a novel framework based on spectral moment representations and constrained optimization, which, under mild moment conditions, achieves the first consistent estimator for such quadratic forms. This approach overcomes fundamental limitations of existing methods and provides a unified treatment for a broad class of high-dimensional statistics. Both theoretical analysis and numerical simulations demonstrate that the proposed method substantially outperforms traditional approaches in applications including optimal portfolio Sharpe ratio estimation and multiple correlation coefficients in regression, effectively resolving the estimation breakdown inherent to the \(p > n\) setting.
This paper addresses high-dimensional inference of a rank-one signal corrupted by sparse graph-structured noise. Specifically, the noise is modeled as the adjacency matrix of a weighted undirected graph with finite average degree. We extend the classical Baik–Ben Arous–Péché (BBP) phase transition to the sparse-graph regime—its first such generalization—by combining the replica method from statistical physics with population dynamics algorithms to solve recursive distributional equations. This yields exact asymptotic characterizations of the largest eigenvalue, the eigenvector density, and the overlap between the signal and the leading eigenvector. We analytically determine the critical signal-to-noise ratio for reliable signal recovery on both Poisson and random regular graphs, and validate our predictions via large-scale numerical diagonalization, observing excellent agreement. Our work establishes the fundamental detection limit of principal component analysis under sparse graph noise and provides a rigorous theoretical foundation and computationally tractable framework for high-dimensional sparse signal inference.
This study addresses the challenges in inferring conditional dependence structures of high-dimensional stationary time series in the frequency domain, which are hindered by truncation and smoothing biases arising from finite-sample discrete Fourier transforms (DFTs) and the difficulty of estimating complex-valued spectral precision matrices in high dimensions. To overcome these issues, the authors propose a debiased complex-valued graphical Lasso estimator based on the full likelihood across neighboring frequency points. The method enables high-dimensional inference of sparse spectral precision matrices at fixed frequencies and achieves entrywise consistent covariance estimation through cross-frequency information aggregation. It constitutes the first approach to perform full-likelihood inference directly on the DFT, effectively controlling regularization, truncation, and smoothing biases. Simulations demonstrate accurate confidence interval coverage across all non-zero frequencies, higher statistical power than existing methods, and false discovery rates close to nominal levels.