🤖 AI Summary
The intrinsic relationship between variable clustering and principal component analysis (PCA) has long been overlooked in the literature.
Method: We propose a novel paradigm that applies K-means clustering to the transpose of the data matrix—thereby clustering variables—and quantifies each cluster’s contribution to individual principal components via variable loadings.
Contribution/Results: This approach establishes, for the first time, an interpretable mapping between variable clusters and the directions of maximal variance in PCA, yielding a unified “variable clustering–PC contribution” analytical framework. Empirical evaluation demonstrates that the method effectively identifies variable groups driving dominant sources of variation, substantially enhancing interpretability in high-dimensional data. It provides a statistically principled yet computationally feasible tool for multivariate exploratory data analysis.
📝 Abstract
Principal Component Analysis (PCA) and K-means constitute fundamental techniques in multivariate analysis. Although they are frequently applied independently or sequentially to cluster observations, the relationship between them, especially when K-means is used to cluster variables rather than observations, has been scarcely explored. This study seeks to address this gap by proposing an innovative method that analyzes the relationship between clusters of variables obtained by applying K-means on transposed data and the principal components of PCA. Our approach involves applying PCA to the original data and K-means to the transposed data set, where the original variables are converted into observations. The contribution of each variable cluster to each principal component is then quantified using measures based on variable loadings. This process provides a tool to explore and understand the clustering of variables and how such clusters contribute to the principal dimensions of variation identified by PCA.