rmt-based model order selection

Designs and analyzes algorithms that determine the number of latent signal components (model order or subspace rank) in multivariate data by deriving and applying random-matrix-theory-based eigenvalue thresholds and bounds to separate signal and noise eigenvalues, and proving their consistency (e.g., almost-sure behavior) in high-dimensional/proportional regimes.

rmt-basedmodelorderselection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Estimating Graph Dimension with Cross-validated Eigenvalues

Aug 06, 2021
FC
Fan Chen
🏛️ University of Wisconsin–Madison

Estimating the latent dimensionality (effective rank) $k$ of graph data is a fundamental challenge in multivariate statistics and network analysis. Existing heuristics—such as the “elbow method”—fail under nonparametric random graph models (e.g., Poisson or Bernoulli edges) due to systematic bias in sample eigenvalues. This paper introduces the first model-agnostic cross-validation framework for $k$: for each sample eigenvector, it conducts an orthogonality hypothesis test against the empirical eigenspace of held-out data, yielding calibrated $p$-values to adaptively identify detectable dimensions. We establish theoretical consistency: under detectability conditions, the estimator converges almost surely to the true $k$, overcoming limitations of ad hoc criteria. Extensive simulations and real-world network analyses demonstrate that our method achieves superior statistical accuracy and computational efficiency compared to classical approaches.

Consistently determining the number of clusters kEstimating latent dimensions in random graph modelsTesting sample eigenvectors' correlation with true dimensions

Testing for latent structure via the Wilcoxon--Wigner random matrix of normalized rank statistics

Dec 21, 2025
JZ
Jonquil Z. Liao
🏛️ University of Wisconsin–Madison

This paper addresses the problem of detecting latent structures—such as communities or principal submatrices—in large symmetric data matrices. We propose a parameter-free, distribution-free, and outlier-robust spectral testing method. Our core methodological innovation is the first systematic construction and analysis of a Wilcoxon–Wigner random matrix framework, which replaces the conventional sample covariance matrix with nonparametric rank-based statistics, thereby eliminating dependence on distributional assumptions or moment conditions. Theoretically, we rigorously establish asymptotic normality for the leading eigenvalue and eigenvector, deriving explicit centering and scaling that yield Gaussian limiting distributions. Practically, the framework enables robust, efficient, and distribution-agnostic hypothesis testing for community detection and principal submatrix localization. This work provides both a novel theoretical foundation and a practical tool for structural inference in high-dimensional symmetric matrices.

Addressing community and principal submatrix detection via spectral methodsDeveloping flexible, efficient, insensitive statistical methodologyTesting latent structure in large symmetric data matrices

Stratified Principal Component Analysis

Jul 28, 2023
TS
Tom Szwagier
🏛️ Université Côte d’Azur | Inria

In principal component analysis (PCA), near-degenerate eigenvalues induce the “isotropy curse,” causing instability in principal direction estimation and diminished interpretability. Method: This paper proposes a novel modeling framework leveraging the eigenvalue multiplicity hierarchy of the covariance matrix. Generalizing probabilistic PCA (PPCA), it employs the ordered geometric structure of flag manifolds to characterize maximum-likelihood estimation under joint eigenvalue multiplicity constraints on signal and noise subspaces. Contribution/Results: We introduce, for the first time, a hierarchical partial-order model selection criterion enabling compact, interpretable low-dimensional modeling. Experiments demonstrate that our method significantly outperforms PPCA—particularly in small-sample regimes and when eigenvalue gaps are weak—achieving superior trade-offs between model complexity and fitting accuracy on both synthetic and real-world data.

Address rotational variability in close-eigenvalue principal componentsPropose simple principal subspace analysis for diverse datasetsTransition from principal components to interpretable principal subspaces

Selecting the number of components in PCA via random signflips.

Dec 05, 2020
DH
D. Hong
🏛️ University of Delaware | University of Pennsylvania

Existing principal component analysis (PCA) model selection methods lack statistical guarantees for determining the number of leading components under heteroscedastic noise—where observation-wise noise variances differ—in high-dimensional settings. Method: We propose Signflip Parallel Analysis (Signflip PA), a novel parallel analysis method that generates an empirical null distribution via random sign flips and adaptively calibrates singular value thresholds. Contribution/Results: Signflip PA is the first to integrate dimension-free operator norm bounds and large-deviation theory for eigenvalues of non-homogeneous matrices into PCA model selection, ensuring consistent factor recovery. We establish its theoretical consistency under a signal-plus-heteroscedastic-noise model. Empirical studies—including simulations and real-data analyses—demonstrate that Signflip PA significantly outperforms classical approaches such as scree plots and conventional parallel analysis, overcoming the fundamental limitation wherein heteroscedasticity causes traditional methods to fail.

Existing methods fail dramatically with varying noise variancesProposes FlipPA for robust rank selection under asymmetric noiseSelecting PCA components lacks guarantees for heterogeneous noise

Spectral Estimators for Structured Generalized Linear Models via Approximate Message Passing

Aug 28, 2023
YZ
Yihan Zhang
🏛️ Institute of Science and Technology Austria | University of Cambridge

Parameter estimation in high-dimensional structured generalized linear models suffers from low efficiency, particularly under realistic design matrices exhibiting anisotropy and strong correlations. Method: This paper introduces a novel spectral estimation framework based on Approximate Message Passing (AMP). Contribution/Results: We provide the first exact asymptotic characterization of spectral estimators under correlated Gaussian designs. We identify a universally optimal covariance-adaptive preprocessing strategy, partially resolving a long-standing conjecture on optimal spectral estimation for rotationally invariant models. Theoretically and empirically, our approach substantially reduces sample complexity and achieves provably statistically optimal estimation accuracy—outperforming existing heuristic methods on canonical designs from computational imaging and genomics.

Characterizing spectral estimators for correlated Gaussian designsEstimating parameters in high-dimensional generalized linear modelsIdentifying optimal preprocessing for efficient parameter estimation

Latest Papers

What's happening recently
View more

PCA recovery thresholds in low-rank matrix inference with sparse noise

Nov 14, 2025
UA
Urte Adomaityte
🏛️ King's College London | University of Bologna

This paper addresses high-dimensional inference of a rank-one signal corrupted by sparse graph-structured noise. Specifically, the noise is modeled as the adjacency matrix of a weighted undirected graph with finite average degree. We extend the classical Baik–Ben Arous–Péché (BBP) phase transition to the sparse-graph regime—its first such generalization—by combining the replica method from statistical physics with population dynamics algorithms to solve recursive distributional equations. This yields exact asymptotic characterizations of the largest eigenvalue, the eigenvector density, and the overlap between the signal and the leading eigenvector. We analytically determine the critical signal-to-noise ratio for reliable signal recovery on both Poisson and random regular graphs, and validate our predictions via large-scale numerical diagonalization, observing excellent agreement. Our work establishes the fundamental detection limit of principal component analysis under sparse graph noise and provides a rigorous theoretical foundation and computationally tractable framework for high-dimensional sparse signal inference.

Analytically identifies critical signal strength for eigenvector-based recoveryGeneralizes BBP transition theory to sparse noise using replica methodsStudies rank-one signal recovery from sparse noise in high-dimensional inference

High-Dimensional Partial Least Squares: Spectral Analysis and Fundamental Limitations

Dec 17, 2025
VL
Victor Léger
🏛️ Université Grenoble Alpes | CNRS | Grenoble INP | GIPSA-lab

This paper investigates the theoretical behavior of high-dimensional partial least squares (PLS) for dual-matrix data fusion, focusing on its ability to estimate and its fundamental limitations in recovering shared low-rank latent structure. Leveraging random matrix theory, we establish the first rigorous asymptotic characterization of PLS-SVD singular vectors’ alignment with true latent directions, quantitatively identifying the phase transition threshold—i.e., the critical signal-to-noise ratio and dimension ratio—governing successful versus failed latent reconstruction. Furthermore, we prove that, for detecting the common latent subspace, PLS-SVD is asymptotically superior to single-dataset PCA, with a theoretically guaranteed advantage. The analysis not only explains the counterintuitive failure of PLS in high dimensions but also precisely delineates its statistical limits and necessary conditions for validity as a multi-view dimensionality reduction method.

Analyzes high-dimensional PLS-SVD's performance in extracting shared latent componentsCharacterizes alignment between estimated and true latent directions asymptoticallyCompares PLS-SVD with separate PCA for common subspace detection

This work addresses the challenging problem of estimating the model order—i.e., the rank of the signal subspace—in the presence of high-dimensional, correlated, and non-Gaussian complex elliptically symmetric (CES) noise. The authors propose a two-stage robust framework: first, a Toeplitz-constrained M-estimator is employed to whiten the unknown covariance structure; second, the signal subspace rank is inferred using large-dimensional random matrix theory (RMT). This approach uniquely integrates Toeplitz-rectified M-estimators—including the sample covariance matrix (SCM), Maronna’s, and Tyler’s estimators—with large-dimensional RMT to construct an almost surely consistent order estimator and derive an explicit eigenvalue separation threshold. Experiments on synthetic data as well as real-world hyperspectral images, electroencephalographic, and financial datasets demonstrate that the proposed method significantly outperforms conventional criteria such as AIC, confirming its robustness and effectiveness.

Complex Elliptically Symmetric noisecorrelated noiselarge-dimensional noise

This study addresses the lack of theoretical justification for principal component analysis (PCA) in high-dimensional factor models when the number of factors is overestimated. It investigates the asymptotic behavior when the true factor number \( r \) is conservatively set to any fixed \( R \geq r \). Leveraging the anisotropic local law from random matrix theory, the paper characterizes the noise-dominated nature, incoherence, and near-orthogonality to true factor loadings of the spurious components. Consistency of factor estimation is established via two rotation mappings, providing the first rigorous theoretical support for the common practice of using a conservative upper bound on the number of factors. The results show that consistent factor estimates are attainable for any fixed \( R \geq r \), enabling \( \sqrt{T} \)-consistent and asymptotically normal inference on treatment effects in factor-augmented regressions.

asymptotic theoryfactor modelhigh-dimensional statistics

Evaluating Singular Value Thresholds for DNN Weight Matrices based on Random Matrix Theory

Dec 14, 2025
KN
Kohei Nishikawa
🏛️ Tokyo University of Science

This paper addresses the lack of theoretical grounding for singular value truncation thresholds in low-rank approximation of deep neural network (DNN) weight matrices. We propose a signal–noise decoupling framework grounded in Random Matrix Theory (RMT), modeling weights as the sum of a low-rank signal component and isotropic random noise, and derive an analytically justified denoising threshold. Furthermore, we introduce—novelty—the first threshold validity metric based on singular vector alignment, quantified as the cosine similarity between the estimated signal’s singular vectors and those of the original weight matrix; this advances beyond conventional empirical threshold selection relying solely on singular value spectra. Experiments across multiple mainstream DNN weight matrices demonstrate that our metric quantitatively distinguishes the signal-preserving capability of competing thresholding methods, leading to significantly improved stability and interpretability of model accuracy after low-rank compression.

Evaluates thresholds for removing singular values in DNN weight matrices.Models weight matrices as signal and noise components.Proposes a metric to assess threshold adequacy using cosine similarity.

Hot Scholars

BG

Banglei Guan

National University of Defense Technology
PhotomechanicsVideometrics
HL

Haifeng Liu

Zhejiang University
Machine LearningData ManagementInformaiton Retrieval
HL

Hongfei Lin

DalianUniversity of Technology
natural language processing,sentimental analysistext miningsocial computing
SJ

S. Joe Qin

Lingnan University, Hong Kong, President, Member of EASA, Fellow of HKAE, NAI, IEEE, IFAC, AIChE
Process data analyticsdata scienceprocess controlsystem identification
AJ

Armando J. Pinho

IEETA/DETI, University of Aveiro
Data CompressionImage CompressionKolmogorov ComplexityAlgorithmic Information Theory