Assessing the impact of dimensionality reduction on clustering performance -- a systematic study

📅 2026-04-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of systematic evaluation regarding how dimensionality reduction methods influence clustering performance. Within a unified framework, it comprehensively assesses the impact of five dimensionality reduction techniques—PCA, Kernel PCA, VAE, Isomap, and MDS—across varying target dimensions on four mainstream clustering algorithms: k-means, Agglomerative Hierarchical Clustering (AHC), Gaussian Mixture Models (GMM), and OPTICS. Clustering quality is quantified using the Adjusted Rand Index (ARI). The work reveals, for the first time, the intricate coupling among intrinsic data geometry, dimensionality reduction strategy, and clustering efficacy, demonstrating that the choice of both reduction method and target dimension must be jointly tailored to the data’s underlying structure and the specific clustering algorithm. Indiscriminate application of dimensionality reduction can substantially degrade clustering performance.

Technology Category

Machine Learning: Dimensionality Reduction/Feature SelectionData Mining & Knowledge Management: Data CompressionSearch and Optimization: Evaluation and Analysis

Application Category

Web Mining and Content Analysis: Normalization, clustering, classification, and summarization of Web textSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsGraph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphs
📝 Abstract
Dimensionality reduction is a critical preprocessing step for clustering high-dimensional data, yet comprehensive evaluation of its impact across diverse methods and data types remains limited. In this study, we systematically assess the influence of five dimensionality reduction techniques - Principal Component Analysis (PCA), Kernel Principal Component Analysis (Kernel PCA), Variational Autoencoder (VAE), Isometric Mapping (Isomap), and Multidimensional Scaling (MDS) - on the performance of four popular clustering algorithms - k-means, Agglomerative Hierarchical Clustering (AHC), Gaussian Mixture Models (GMM), and Ordering Points to Identify the Clustering Structure (OPTICS). We evaluate clustering quality using the Adjusted Rand Index (ARI), comparing results without and with dimensionality reduction at different reduction levels recommended in the literature (i.e., k-1, where k is the number of clusters, and 25% and 50% of the original number of dimensions). Our findings underscore the importance of a careful selection of the dimensionality reduction technique and the dimensionality reduction level that should be tailored to intrinsic data geometry and clustering algorithms under consideration.
Problem

Research questions and friction points this paper is trying to address.

dimensionality reduction
clustering performance
high-dimensional data
systematic evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

dimensionality reduction
clustering performance
systematic evaluation
Adjusted Rand Index
data geometry
🔎 Similar Papers
2021-06-14IEEE Transactions on Visualization and Computer GraphicsCitations: 12
💼 Related Jobs
No related jobs found.
O
Ousmane Assani Amate
Université du Québec à Montréal, Montreal, Quebec, Canada
Mohammadreza Bakhtyari
Mohammadreza Bakhtyari
Ph.D. student
Machine LearningDeep Learning
É
Émilie Roy
Université du Québec à Montréal, Montreal, Quebec, Canada
Vladimir Makarenkov
Vladimir Makarenkov
Université du Québec à Montréal
bioinformaticsdata miningclusteringmachine learningoperations research