An Approach to Variable Clustering: K-means in Transposed Data and its Relationship with Principal Component Analysis

📅 2025-11-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
The intrinsic relationship between variable clustering and principal component analysis (PCA) has long been overlooked in the literature. Method: We propose a novel paradigm that applies K-means clustering to the transpose of the data matrix—thereby clustering variables—and quantifies each cluster’s contribution to individual principal components via variable loadings. Contribution/Results: This approach establishes, for the first time, an interpretable mapping between variable clusters and the directions of maximal variance in PCA, yielding a unified “variable clustering–PC contribution” analytical framework. Empirical evaluation demonstrates that the method effectively identifies variable groups driving dominant sources of variation, substantially enhancing interpretability in high-dimensional data. It provides a statistically principled yet computationally feasible tool for multivariate exploratory data analysis.

Technology Category

Machine Learning: ClusteringComputer Vision: Interpretability, Explainability, and TransparencyData Mining & Knowledge Management: Data Compression

Application Category

Web Mining and Content Analysis: Normalization, clustering, classification, and summarization of Web textGraph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphsUser Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalization
📝 Abstract
Principal Component Analysis (PCA) and K-means constitute fundamental techniques in multivariate analysis. Although they are frequently applied independently or sequentially to cluster observations, the relationship between them, especially when K-means is used to cluster variables rather than observations, has been scarcely explored. This study seeks to address this gap by proposing an innovative method that analyzes the relationship between clusters of variables obtained by applying K-means on transposed data and the principal components of PCA. Our approach involves applying PCA to the original data and K-means to the transposed data set, where the original variables are converted into observations. The contribution of each variable cluster to each principal component is then quantified using measures based on variable loadings. This process provides a tool to explore and understand the clustering of variables and how such clusters contribute to the principal dimensions of variation identified by PCA.
Problem

Research questions and friction points this paper is trying to address.

Explores relationship between variable clustering via transposed K-means and PCA components.
Quantifies variable clusters' contributions to principal components using loading-based measures.
Provides tool to understand variable clustering and its impact on PCA variation dimensions.
Innovation

Methods, ideas, or system contributions that make the work stand out.

K-means clustering applied to transposed data
Quantifying variable clusters' contributions to principal components
Linking variable clustering with PCA dimensions
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
University of Cuenca | UNEMI | Universidad Estatal de Milagro
Victor Saquicela
Victor Saquicela
Department of Computer Science, University of Cuenca, Cuenca, Ecuador
K
Kenneth Palacio Baus
Department of Electrical Engineering, University of Cuenca, Cuenca, Ecuador
M
Mario Chifla
UNEMI, Universidad Estatal de Milagro, Milagro, Ecuador