🤖 AI Summary
This paper investigates the theoretical behavior of high-dimensional partial least squares (PLS) for dual-matrix data fusion, focusing on its ability to estimate and its fundamental limitations in recovering shared low-rank latent structure. Leveraging random matrix theory, we establish the first rigorous asymptotic characterization of PLS-SVD singular vectors’ alignment with true latent directions, quantitatively identifying the phase transition threshold—i.e., the critical signal-to-noise ratio and dimension ratio—governing successful versus failed latent reconstruction. Furthermore, we prove that, for detecting the common latent subspace, PLS-SVD is asymptotically superior to single-dataset PCA, with a theoretically guaranteed advantage. The analysis not only explains the counterintuitive failure of PLS in high dimensions but also precisely delineates its statistical limits and necessary conditions for validity as a multi-view dimensionality reduction method.
📝 Abstract
Partial Least Squares (PLS) is a widely used method for data integration, designed to extract latent components shared across paired high-dimensional datasets. Despite decades of practical success, a precise theoretical understanding of its behavior in high-dimensional regimes remains limited. In this paper, we study a data integration model in which two high-dimensional data matrices share a low-rank common latent structure while also containing individual-specific components. We analyze the singular vectors of the associated cross-covariance matrix using tools from random matrix theory and derive asymptotic characterizations of the alignment between estimated and true latent directions. These results provide a quantitative explanation of the reconstruction performance of the PLS variant based on Singular Value Decomposition (PLS-SVD) and identify regimes where the method exhibits counter-intuitive or limiting behavior. Building on this analysis, we compare PLS-SVD with principal component analysis applied separately to each dataset and show its asymptotic superiority in detecting the common latent subspace. Overall, our results offer a comprehensive theoretical understanding of high-dimensional PLS-SVD, clarifying both its advantages and fundamental limitations.