Score
Designs and computes a linear matrix (the whitening matrix) that decorrelates multivariate data and rescales feature variances so the transformed data have identity covariance, and implements or applies that transformation to feature vectors or data matrices. Analyzes the whitening operator’s spectral properties (via SVD/EVD), chooses or constructs whitening procedures (ZCA, PCA-whitening, etc.), and evaluates how whitening interacts with noise models such as isotropic noise.
This study addresses the vulnerability of linear probes to spurious correlations, which arises from their inherent bias toward directions associated with large eigenvalues and consequently degrades generalization performance. To mitigate this issue, the paper establishes a theoretical connection between linear probing and maximum-margin classifiers, proposing a covariance whitening strategy that equalizes the eigenvalue spectrum and thereby eliminates directional biases. Notably, this approach operates as a general-purpose preprocessing technique that requires neither prior knowledge nor labeled data. Experiments on both synthetic datasets and standard benchmarks demonstrate that the proposed whitening strategy significantly enhances the robustness of linear probes against spurious correlations while effectively improving the performance of existing methods.
In high-dimensional sparse regimes (large dimension-to-sample ratio), spectral distortion of the sample covariance matrix renders standard whitening ineffective: after whitening, the means of a spherical Gaussian mixture model (GMM) fail to become asymptotically orthogonal, severely degrading latent variable estimation based on higher-order moment tensor decomposition. This paper establishes, for the first time, a rigorous characterization—grounded in random matrix theory—of the asymptotic behavior of inner products between whitened means in high dimensions. Leveraging this analysis, we propose a spectrally corrected whitening matrix that achieves asymptotic orthogonality of the means. This correction restores the decomposability of higher-order moment tensors and significantly improves GMM parameter estimation accuracy. Empirical results demonstrate that the proposed corrected whitening remains robust and effective even in finite-sample settings, thereby overcoming fundamental theoretical and practical limitations of conventional whitening in large-dimensional sparse scenarios.
In principal component analysis (PCA), near-degenerate eigenvalues induce the “isotropy curse,” causing instability in principal direction estimation and diminished interpretability. Method: This paper proposes a novel modeling framework leveraging the eigenvalue multiplicity hierarchy of the covariance matrix. Generalizing probabilistic PCA (PPCA), it employs the ordered geometric structure of flag manifolds to characterize maximum-likelihood estimation under joint eigenvalue multiplicity constraints on signal and noise subspaces. Contribution/Results: We introduce, for the first time, a hierarchical partial-order model selection criterion enabling compact, interpretable low-dimensional modeling. Experiments demonstrate that our method significantly outperforms PPCA—particularly in small-sample regimes and when eigenvalue gaps are weak—achieving superior trade-offs between model complexity and fitting accuracy on both synthetic and real-world data.
Existing asymmetric kernel singular value decomposition (KSVD) methods rely on finite-dimensional approximations, limiting their ability to handle infinite-dimensional feature maps, and their variational objectives may be unbounded. Method: We propose the Coupled Covariance Eigenproblem (CCE) framework—the first rigorous variational formulation of KSVD in infinite-dimensional Hilbert spaces—unifying asymmetric KSVD with covariance operator theory and accommodating arbitrary non-Mercer, asymmetric kernels. We further derive an asymmetric Nyström method based on coupled adjoint eigenfunctions, overcoming classical limitations of symmetric kernel approximation or linear SVD modeling. Contribution/Results: Experiments demonstrate that our method significantly outperforms symmetric-baseline and linear-SVD approaches across multiple tasks, achieving faster training convergence and improved generalization. This work provides the first empirical validation of the practical utility of asymmetric kernel learning.
Parameter estimation in high-dimensional structured generalized linear models suffers from low efficiency, particularly under realistic design matrices exhibiting anisotropy and strong correlations. Method: This paper introduces a novel spectral estimation framework based on Approximate Message Passing (AMP). Contribution/Results: We provide the first exact asymptotic characterization of spectral estimators under correlated Gaussian designs. We identify a universally optimal covariance-adaptive preprocessing strategy, partially resolving a long-standing conjecture on optimal spectral estimation for rotationally invariant models. Theoretically and empirically, our approach substantially reduces sample complexity and achieves provably statistically optimal estimation accuracy—outperforming existing heuristic methods on canonical designs from computational imaging and genomics.
This study addresses the distortion of scientific interpretation in principal component analysis (PCA) biplots caused by standardization and projection practices. It proposes three formal conditions for interpretive validity—objective alignment, spectral identifiability, and projection sufficiency—and, for the first time, explicitly distinguishes scientific interpretability from computational correctness. The authors introduce a basis-invariant diagnostic framework via projection operators and quantify pairwise errors induced by omitted coordinates using residual Gram bounds, thereby establishing a verifiable approach for diagnosis and correction. The method’s efficacy is demonstrated across six controlled population scenarios and real-data case studies, revealing common interpretive pitfalls and supporting Bootstrap-based extensions.
本文针对高维数据的主成分分析问题,提出了一种基于子空间分析的新算法SuperPCA,通过利用样本协方差矩阵的部分特征空间来更高效准确地估计主要信号。
研究探讨了白化嵌入的平方范数作为无训练似然替代的问题,通过分析白化逆转编码器谱层次结构机制,指出其更适合解释为马氏距离而非对数似然。
研究通过分析自编码器参数矩阵的谱特性,将其视为数据样本的向量表示,从而区分不同训练子集的模型,无需使用原始样本。
This study addresses the frequent misinterpretation of fluctuations in dominant eigensubspaces and scalar spectral functionals—such as the absorption ratio—in rolling covariance estimation as genuine market structural changes, when they are often artifacts of estimation noise, particularly under shrinkage. By leveraging perturbation analysis and calibrated inference, the work derives, for the first time, the first-order null distribution of eigensubspace variation under overlapping windows and establishes its invariance under rotation-equivariant shrinkage estimators. It further shows that only scale-invariant spectral functionals enjoy first-order immunity to elliptical kurtosis. To correct high-dimensional bias in the absorption ratio, a trace-preserving spiked debiased estimator is proposed. Theoretical results, supported by Davis–Kahan bounds, distribution-free confidence bands, and an estimator-aware bootstrap, are validated through simulations and successfully applied to equity data for reliable detection of true market structural shifts.