Score
Design and implement hybrid principal component estimators that produce stable estimates of leading principal eigenvectors and eigenvalues from multivariate data/covariance matrices by combining estimation strategies (e.g., shrinkage, projection, or regularization) to reduce sensitivity to sampling variability and near-eigenvalue multiplicity. Analyze and validate these estimators by proving consistency and asymptotic normality, quantifying robustness to close or ill-conditioned eigenvalue spectra, and integrating the procedures into downstream multivariate analyses.
Conventional PCA fails for max-stable distributions due to their restricted support and heavy tails, which violate PCA’s reliance on finite second moments and linear variance structure. Method: We propose the first PCA paradigm tailored to extreme-value statistics, based on a max-linear regression framework that preserves max-stability in low-dimensional projections. Our approach integrates extremal regression, nonlinear projection optimization, and stability-constrained inference. Contribution/Results: Theoretically, we establish necessary and sufficient conditions for perfect reconstruction and prove consistent estimability of the optimal projection matrix. Empirically, simulations and real-world extreme-value data demonstrate substantial improvements in low-dimensional representation fidelity and reconstruction accuracy for heavy-tailed extremes. The method provides an interpretable, computationally tractable dimensionality reduction foundation for high-dimensional extreme-value modeling.
In principal component analysis (PCA), near-degenerate eigenvalues induce the “isotropy curse,” causing instability in principal direction estimation and diminished interpretability. Method: This paper proposes a novel modeling framework leveraging the eigenvalue multiplicity hierarchy of the covariance matrix. Generalizing probabilistic PCA (PPCA), it employs the ordered geometric structure of flag manifolds to characterize maximum-likelihood estimation under joint eigenvalue multiplicity constraints on signal and noise subspaces. Contribution/Results: We introduce, for the first time, a hierarchical partial-order model selection criterion enabling compact, interpretable low-dimensional modeling. Experiments demonstrate that our method significantly outperforms PPCA—particularly in small-sample regimes and when eigenvalue gaps are weak—achieving superior trade-offs between model complexity and fitting accuracy on both synthetic and real-world data.
Existing spectral shrinkage methods for high-dimensional covariance estimation rely either on heuristic objectives or strong distributional assumptions. To address this, we propose a geometrically unified framework grounded in distributionally robust optimization (DRO). Our approach defines geometric divergence measures—including KL, Fisher–Rao, and Wasserstein divergences—on the covariance manifold, constructs data-driven ambiguity sets, and automatically derives spectral shrinkage estimators by minimizing the worst-case Frobenius error. Crucially, we establish for the first time that the functional form of shrinkage is intrinsically determined by the geometric structure of the chosen divergence, thereby eliminating dependence on prior distributional assumptions. We prove asymptotic consistency of the proposed estimator and derive an $O(1/n)$ finite-sample error bound. Moreover, we obtain three explicit robust estimators, which match state-of-the-art performance on both synthetic and real-world datasets.
Parameter estimation in high-dimensional structured generalized linear models suffers from low efficiency, particularly under realistic design matrices exhibiting anisotropy and strong correlations. Method: This paper introduces a novel spectral estimation framework based on Approximate Message Passing (AMP). Contribution/Results: We provide the first exact asymptotic characterization of spectral estimators under correlated Gaussian designs. We identify a universally optimal covariance-adaptive preprocessing strategy, partially resolving a long-standing conjecture on optimal spectral estimation for rotationally invariant models. Theoretically and empirically, our approach substantially reduces sample complexity and achieves provably statistically optimal estimation accuracy—outperforming existing heuristic methods on canonical designs from computational imaging and genomics.
Estimating the latent dimensionality (effective rank) $k$ of graph data is a fundamental challenge in multivariate statistics and network analysis. Existing heuristics—such as the “elbow method”—fail under nonparametric random graph models (e.g., Poisson or Bernoulli edges) due to systematic bias in sample eigenvalues. This paper introduces the first model-agnostic cross-validation framework for $k$: for each sample eigenvector, it conducts an orthogonality hypothesis test against the empirical eigenspace of held-out data, yielding calibrated $p$-values to adaptively identify detectable dimensions. We establish theoretical consistency: under detectability conditions, the estimator converges almost surely to the true $k$, overcoming limitations of ad hoc criteria. Extensive simulations and real-world network analyses demonstrate that our method achieves superior statistical accuracy and computational efficiency compared to classical approaches.
该研究解决了高维因子模型中主成分估计误差问题,通过将误差分解为可解释的两部分,并在金融经济、基因组学等领域应用。
This study addresses the frequent misinterpretation of fluctuations in dominant eigensubspaces and scalar spectral functionals—such as the absorption ratio—in rolling covariance estimation as genuine market structural changes, when they are often artifacts of estimation noise, particularly under shrinkage. By leveraging perturbation analysis and calibrated inference, the work derives, for the first time, the first-order null distribution of eigensubspace variation under overlapping windows and establishes its invariance under rotation-equivariant shrinkage estimators. It further shows that only scale-invariant spectral functionals enjoy first-order immunity to elliptical kurtosis. To correct high-dimensional bias in the absorption ratio, a trace-preserving spiked debiased estimator is proposed. Theoretical results, supported by Davis–Kahan bounds, distribution-free confidence bands, and an estimator-aware bootstrap, are validated through simulations and successfully applied to equity data for reliable detection of true market structural shifts.
This study investigates the capacity and stability of principal component analysis (PCA) to recover latent structures at observational scales ranging from billions to trillions. Through large-scale empirical experiments on datasets with 10 billion and 1 trillion samples—both random and containing embedded latent factors—this work provides the first validation of PCA’s convergence behavior at the trillion-sample scale. The results demonstrate that PCA achieves practical convergence well before reaching a trillion observations: outputs remain highly stable on random data, while in datasets with latent factors, the top three principal components account for 99.996% of the variance, accurately reconstructing the underlying ground-truth structure. These findings offer both theoretical grounding and empirical evidence supporting the application of PCA to ultra-large-scale data.
本文针对高维数据的主成分分析问题,提出了一种基于子空间分析的新算法SuperPCA,通过利用样本协方差矩阵的部分特征空间来更高效准确地估计主要信号。
This study addresses the insufficient estimation accuracy of principal eigenvectors under high-dimensional proportional asymptotics by proposing a data-adaptive augmented James–Stein shrinkage framework. Built upon the generalized spiked model and subspace augmentation techniques, the proposed method incorporates auxiliary information to optimize eigenspace estimation. Notably, it automatically reduces to principal component analysis (PCA) when such auxiliary information is unavailable, thereby ensuring both flexibility and robustness. Theoretical analyses and simulation studies demonstrate that this framework strictly outperforms PCA and existing high-dimension, low-sample-size (HDLSS) methods, yielding substantial improvements in finite-sample settings. Furthermore, the approach exhibits strong robustness against the misspecification of the number of spikes.