Score
Designs and analyzes algorithms that determine the number of latent signal components (model order or subspace rank) in multivariate data by deriving and applying random-matrix-theory-based eigenvalue thresholds and bounds to separate signal and noise eigenvalues, and proving their consistency (e.g., almost-sure behavior) in high-dimensional/proportional regimes.
Estimating the latent dimensionality (effective rank) $k$ of graph data is a fundamental challenge in multivariate statistics and network analysis. Existing heuristics—such as the “elbow method”—fail under nonparametric random graph models (e.g., Poisson or Bernoulli edges) due to systematic bias in sample eigenvalues. This paper introduces the first model-agnostic cross-validation framework for $k$: for each sample eigenvector, it conducts an orthogonality hypothesis test against the empirical eigenspace of held-out data, yielding calibrated $p$-values to adaptively identify detectable dimensions. We establish theoretical consistency: under detectability conditions, the estimator converges almost surely to the true $k$, overcoming limitations of ad hoc criteria. Extensive simulations and real-world network analyses demonstrate that our method achieves superior statistical accuracy and computational efficiency compared to classical approaches.
This paper addresses the problem of detecting latent structures—such as communities or principal submatrices—in large symmetric data matrices. We propose a parameter-free, distribution-free, and outlier-robust spectral testing method. Our core methodological innovation is the first systematic construction and analysis of a Wilcoxon–Wigner random matrix framework, which replaces the conventional sample covariance matrix with nonparametric rank-based statistics, thereby eliminating dependence on distributional assumptions or moment conditions. Theoretically, we rigorously establish asymptotic normality for the leading eigenvalue and eigenvector, deriving explicit centering and scaling that yield Gaussian limiting distributions. Practically, the framework enables robust, efficient, and distribution-agnostic hypothesis testing for community detection and principal submatrix localization. This work provides both a novel theoretical foundation and a practical tool for structural inference in high-dimensional symmetric matrices.
In principal component analysis (PCA), near-degenerate eigenvalues induce the “isotropy curse,” causing instability in principal direction estimation and diminished interpretability. Method: This paper proposes a novel modeling framework leveraging the eigenvalue multiplicity hierarchy of the covariance matrix. Generalizing probabilistic PCA (PPCA), it employs the ordered geometric structure of flag manifolds to characterize maximum-likelihood estimation under joint eigenvalue multiplicity constraints on signal and noise subspaces. Contribution/Results: We introduce, for the first time, a hierarchical partial-order model selection criterion enabling compact, interpretable low-dimensional modeling. Experiments demonstrate that our method significantly outperforms PPCA—particularly in small-sample regimes and when eigenvalue gaps are weak—achieving superior trade-offs between model complexity and fitting accuracy on both synthetic and real-world data.
Existing principal component analysis (PCA) model selection methods lack statistical guarantees for determining the number of leading components under heteroscedastic noise—where observation-wise noise variances differ—in high-dimensional settings. Method: We propose Signflip Parallel Analysis (Signflip PA), a novel parallel analysis method that generates an empirical null distribution via random sign flips and adaptively calibrates singular value thresholds. Contribution/Results: Signflip PA is the first to integrate dimension-free operator norm bounds and large-deviation theory for eigenvalues of non-homogeneous matrices into PCA model selection, ensuring consistent factor recovery. We establish its theoretical consistency under a signal-plus-heteroscedastic-noise model. Empirical studies—including simulations and real-data analyses—demonstrate that Signflip PA significantly outperforms classical approaches such as scree plots and conventional parallel analysis, overcoming the fundamental limitation wherein heteroscedasticity causes traditional methods to fail.
Parameter estimation in high-dimensional structured generalized linear models suffers from low efficiency, particularly under realistic design matrices exhibiting anisotropy and strong correlations. Method: This paper introduces a novel spectral estimation framework based on Approximate Message Passing (AMP). Contribution/Results: We provide the first exact asymptotic characterization of spectral estimators under correlated Gaussian designs. We identify a universally optimal covariance-adaptive preprocessing strategy, partially resolving a long-standing conjecture on optimal spectral estimation for rotationally invariant models. Theoretically and empirically, our approach substantially reduces sample complexity and achieves provably statistically optimal estimation accuracy—outperforming existing heuristic methods on canonical designs from computational imaging and genomics.
This paper addresses high-dimensional inference of a rank-one signal corrupted by sparse graph-structured noise. Specifically, the noise is modeled as the adjacency matrix of a weighted undirected graph with finite average degree. We extend the classical Baik–Ben Arous–Péché (BBP) phase transition to the sparse-graph regime—its first such generalization—by combining the replica method from statistical physics with population dynamics algorithms to solve recursive distributional equations. This yields exact asymptotic characterizations of the largest eigenvalue, the eigenvector density, and the overlap between the signal and the leading eigenvector. We analytically determine the critical signal-to-noise ratio for reliable signal recovery on both Poisson and random regular graphs, and validate our predictions via large-scale numerical diagonalization, observing excellent agreement. Our work establishes the fundamental detection limit of principal component analysis under sparse graph noise and provides a rigorous theoretical foundation and computationally tractable framework for high-dimensional sparse signal inference.
This paper investigates the theoretical behavior of high-dimensional partial least squares (PLS) for dual-matrix data fusion, focusing on its ability to estimate and its fundamental limitations in recovering shared low-rank latent structure. Leveraging random matrix theory, we establish the first rigorous asymptotic characterization of PLS-SVD singular vectors’ alignment with true latent directions, quantitatively identifying the phase transition threshold—i.e., the critical signal-to-noise ratio and dimension ratio—governing successful versus failed latent reconstruction. Furthermore, we prove that, for detecting the common latent subspace, PLS-SVD is asymptotically superior to single-dataset PCA, with a theoretically guaranteed advantage. The analysis not only explains the counterintuitive failure of PLS in high dimensions but also precisely delineates its statistical limits and necessary conditions for validity as a multi-view dimensionality reduction method.
This work addresses the challenging problem of estimating the model order—i.e., the rank of the signal subspace—in the presence of high-dimensional, correlated, and non-Gaussian complex elliptically symmetric (CES) noise. The authors propose a two-stage robust framework: first, a Toeplitz-constrained M-estimator is employed to whiten the unknown covariance structure; second, the signal subspace rank is inferred using large-dimensional random matrix theory (RMT). This approach uniquely integrates Toeplitz-rectified M-estimators—including the sample covariance matrix (SCM), Maronna’s, and Tyler’s estimators—with large-dimensional RMT to construct an almost surely consistent order estimator and derive an explicit eigenvalue separation threshold. Experiments on synthetic data as well as real-world hyperspectral images, electroencephalographic, and financial datasets demonstrate that the proposed method significantly outperforms conventional criteria such as AIC, confirming its robustness and effectiveness.
This study addresses the lack of theoretical justification for principal component analysis (PCA) in high-dimensional factor models when the number of factors is overestimated. It investigates the asymptotic behavior when the true factor number \( r \) is conservatively set to any fixed \( R \geq r \). Leveraging the anisotropic local law from random matrix theory, the paper characterizes the noise-dominated nature, incoherence, and near-orthogonality to true factor loadings of the spurious components. Consistency of factor estimation is established via two rotation mappings, providing the first rigorous theoretical support for the common practice of using a conservative upper bound on the number of factors. The results show that consistent factor estimates are attainable for any fixed \( R \geq r \), enabling \( \sqrt{T} \)-consistent and asymptotically normal inference on treatment effects in factor-augmented regressions.
This paper addresses the lack of theoretical grounding for singular value truncation thresholds in low-rank approximation of deep neural network (DNN) weight matrices. We propose a signal–noise decoupling framework grounded in Random Matrix Theory (RMT), modeling weights as the sum of a low-rank signal component and isotropic random noise, and derive an analytically justified denoising threshold. Furthermore, we introduce—novelty—the first threshold validity metric based on singular vector alignment, quantified as the cosine similarity between the estimated signal’s singular vectors and those of the original weight matrix; this advances beyond conventional empirical threshold selection relying solely on singular value spectra. Experiments across multiple mainstream DNN weight matrices demonstrate that our metric quantitatively distinguishes the signal-preserving capability of competing thresholding methods, leading to significantly improved stability and interpretability of model accuracy after low-rank compression.