🤖 AI Summary
Selecting appropriate clustering algorithms for unsupervised learning remains a longstanding challenge due to the absence of ground-truth labels and the high sensitivity of algorithm performance to data characteristics.
Method: This paper introduces a large-scale benchmark comprising 34,000 synthetic datasets and proposes an end-to-end trainable deep neural architecture that integrates convolutional layers, residual connections, and self-attention mechanisms to directly learn algorithm suitability from raw data—eliminating reliance on handcrafted meta-features or conventional clustering validity indices (CVIs). The model jointly captures local and global structural patterns and is supervised using Adjusted Rand Index (ARI) scores across ten mainstream clustering algorithms.
Contribution/Results: On synthetic data, the method achieves an ARI improvement of 0.497; on real-world benchmarks, it outperforms the best AutoML-based approach by 15.3% in accuracy, significantly surpassing existing CVIs and automated clustering selection methods.
📝 Abstract
We introduce ClustRecNet - a novel deep learning (DL)-based recommendation framework for determining the most suitable clustering algorithms for a given dataset, addressing the long-standing challenge of clustering algorithm selection in unsupervised learning. To enable supervised learning in this context, we construct a comprehensive data repository comprising 34,000 synthetic datasets with diverse structural properties. Each of them was processed using 10 popular clustering algorithms. The resulting clusterings were assessed via the Adjusted Rand Index (ARI) to establish ground truth labels, used for training and evaluation of our DL model. The proposed network architecture integrates convolutional, residual, and attention mechanisms to capture both local and global structural patterns from the input data. This design supports end-to-end training to learn compact representations of datasets and enables direct recommendation of the most suitable clustering algorithm, reducing reliance on handcrafted meta-features and traditional Cluster Validity Indices (CVIs). Comprehensive experiments across synthetic and real-world benchmarks demonstrate that our DL model consistently outperforms conventional CVIs (e.g. Silhouette, Calinski-Harabasz, Davies-Bouldin, and Dunn) as well as state-of-the-art AutoML clustering recommendation approaches (e.g. ML2DAC, AutoCluster, and AutoML4Clust). Notably, the proposed model achieves a 0.497 ARI improvement over the Calinski-Harabasz index on synthetic data and a 15.3% ARI gain over the best-performing AutoML approach on real-world data.