🤖 AI Summary
This study addresses the failure of Parallel Analysis (PA) in high-dimensional data caused by mixed signals and heavy-tailed distributions, which complicates latent dimension identification. To overcome this limitation, we propose Rank PA, a method that restores spectral exchangeability via rank transformation. This approach enables accurate nonparametric estimation of the number of latent factors without assuming factor pervasiveness or imposing higher-order moment constraints. Theoretically, consistency is established under relaxed conditions, achieving factor number estimation at an $O(n^{1/3-\tau})$ rate. Empirically, the method demonstrates robust performance on both simulated data and the FRED-MD dataset. By transcending the limitations of traditional PA, this work provides a reliable nonparametric inference framework for high-dimensional heavy-tailed data.
📝 Abstract
Principal component analysis and factor analysis are foundational to the study of high-dimensional data, yet their efficacy depends entirely on correctly identifying the latent dimension $r$. While parallel analysis (PA) is widely regarded as the gold standard for this task, its performance suffers in high-dimensional regimes consisting of mixed signals and/or heavy-tailed error distributions. This arises because under the preceding conditions, column-wise permutation mechanism (or any modern variations) applied to the data matrix $\textbf{X}$ fails to destroy the signal, leading to an inflated null spectrum. To address this, we propose a nonparametric solution. We show that the latent dimension is uniquely preserved in the degree of order within the rank-transformed data matrix, which restores spectral exchangeability under permutation. We prove the consistency of Rank PA under relaxed conditions, requiring neither the pervasiveness of factors nor restrictive fourth moment assumptions. We establish that Rank PA can consistently estimate the number of factors up to the maximum of $O(n^{1/3 - τ})$ where $τ>0$. The robust performance of our remedy is demonstrated through numerical simulations and an analysis of the FRED-MD dataset.