🤖 AI Summary
This study systematically evaluates the effectiveness of UMAP and its supervised variants in leveraging label information for dimensionality reduction in both regression and classification tasks, presenting the first in-depth analysis of supervised UMAP in regression settings. Through comprehensive comparisons across synthetic and real-world datasets—including UMAP, supervised UMAP, PCA, Kernel PCA, Sliced Inverse Regression (SIR), Kernel SIR, and t-SNE—the quality of low-dimensional embeddings is assessed by their predictive performance. The findings reveal that while supervised UMAP excels in classification tasks, it struggles to effectively incorporate response variable information in regression scenarios, leading to substantially degraded performance. This limitation underscores a critical methodological gap in current supervised UMAP formulations for regression and provides clear guidance for future algorithmic improvements.
📝 Abstract
Uniform Manifold Approximation and Projection (UMAP) is a widely used manifold learning technique for dimensionality reduction. This paper studies UMAP, supervised UMAP, and several competing dimensionality reduction methods, including Principal Component Analysis (PCA), Kernel PCA, Sliced Inverse Regression (SIR), Kernel SIR, and t-distributed Stochastic Neighbor Embedding, through a comprehensive comparative analysis. Although UMAP has attracted substantial attention for preserving local and global structures, its supervised extensions, particularly for regression settings, remain rather underexplored. We provide a systematic evaluation of supervised UMAP for both regression and classification using simulated and real datasets, with performance assessed via predictive accuracy on low-dimensional embeddings. Our results show that supervised UMAP performs well for classification but exhibits limitations in effectively incorporating response information for regression, highlighting an important direction for future development.