Score
Designs and implements end-to-end clustering pipelines that group items, records, documents, or users by semantic or feature similarity — selecting and configuring unsupervised algorithms (hierarchical, density‑based, embedding/representation‑based, and other scalable clustering methods), applying dimensionality reduction or embedding techniques, and tuning for large datasets. Builds evaluation and interpretation processes to assign interpretable cluster memberships or labels, link clusters to dominant features or themes, and compare clustering outcomes with classification approaches for analysis.
This paper addresses the challenge of evaluating and selecting appropriate clustering algorithms—K-means, DBSCAN, and spectral clustering—for high-dimensional data. We propose a unified unsupervised evaluation framework that integrates multiple dimensionality reduction techniques (PCA, t-SNE, UMAP) with composite validity metrics (silhouette coefficient, Adjusted Rand Index, Normalized Mutual Information). For the first time, we systematically uncover synergistic interactions between dimensionality reduction methods and clustering algorithms: UMAP preprocessing significantly enhances spectral clustering performance on complex manifold-structured data (e.g., MNIST, Fashion-MNIST, UCI HAR); K-means maintains superior computational efficiency; and DBSCAN excels at detecting irregularly shaped clusters. The framework provides reproducible, interpretable, and empirically grounded guidance for algorithm selection in high-dimensional clustering tasks.
Existing clustering approaches suffer from paradigm fragmentation, ambiguous applicability, and limited interpretability when handling high-dimensional, dynamic, and heterogeneous data. Method: This study systematically unifies five major clustering paradigms—partitioning, hierarchical, data-stream, subspace, and network clustering—and establishes, for the first time, a cross-paradigm applicability mapping framework that rigorously defines their extension boundaries and adaptation conditions in dynamic data streams and heterogeneous networks. A comprehensive evaluation of representative algorithms (K-means, DBSCAN, AGNES, BIRCH, CLIQUE, and community detection methods) is conducted using silhouette coefficient, F1-score, and visualization analysis. Contribution/Results: We propose a reusable three-stage practical guideline—preprocessing, algorithm selection, and validation—and deliver an application atlas spanning 12 disciplines. The framework significantly improves clustering accuracy and interpretability on high-dimensional sparse data.
Existing centroid-based objectives (e.g., k-means) lack efficient and exact optimization over arbitrary valid hierarchical clusterings, hindering flexible, interpretable hierarchical modeling. Method: We propose a tree-structured optimization framework grounded in ultrametrics and dynamic programming. Crucially, we prove that—given any initial clustering tree—the optimal hierarchical solution for centroid-based objectives (e.g., k-means) at all granularities can be computed exactly in *O*(*nk*) time. The resulting hierarchy inherently satisfies hierarchical consistency and enables rapid enumeration of semantically equivalent hierarchical structures from the same base tree. Results: Extensive experiments across multiple datasets demonstrate substantial improvements in flexibility, computational efficiency, and generalization performance of hierarchical clustering. Our framework establishes a new paradigm unifying exact objective optimization with interpretable, semantics-aware hierarchical representation learning.
High-stakes domains—such as healthcare and finance—demand interpretable clustering outcomes to ensure transparency, accountability, and regulatory compliance. Method: This survey systematically analyzes over 120 scholarly works, proposing the first unified taxonomy of interpretability dimensions for clustering. It rigorously distinguishes intrinsically interpretable models—including rule-based, prototype-based, and sparsity-driven approaches—from post-hoc explanation techniques—such as visualization, feature attribution, and local surrogate modeling. The study further develops a use-case-oriented, structured classification framework and principled evaluation criteria. Contribution/Results: It introduces the first practical guideline for selecting appropriate interpretable clustering methods based on application requirements. The work bridges theoretical foundations with real-world deployment, providing both conceptual clarity and actionable insights to support the development and adoption of clustering algorithms that jointly optimize accuracy and interpretability—thereby advancing trustworthy AI in ethically and regulatorily sensitive contexts.
Selecting optimal clustering algorithms for large-scale data remains challenging, particularly under semi-supervised settings where only costly oracle queries yield partial ground-truth labels. Method: This paper introduces and formalizes the novel “scale generalization” theory: if an algorithm is optimal on a small subsample, it remains optimal on the full dataset. We derive sufficient conditions for scale generalization for single-linkage, k-means++, and smoothed Gonzalez k-center algorithms. Our framework integrates clustering stability analysis, sampling theory, and semi-supervised learning, employing subsampling-based evaluation coupled with oracle validation. Results: Empirical evaluation across multiple real-world datasets demonstrates that just 5% of the data suffices to identify the globally optimal algorithm with 100% accuracy. The core contribution is the first provably sound and empirically verifiable theory of scale generalization for clustering algorithms, enabling efficient, robust algorithm selection without full-label supervision.
This work addresses the limitations of traditional hierarchical clustering, which relies solely on pairwise distances and struggles to capture density variations and local connectivity inherent in graph-structured data. The authors propose a novel integrated approach that synergistically combines hierarchical, density-based, and graph clustering paradigms. Specifically, they construct KNN-induced subgraphs and introduce a cluster similarity measure that jointly accounts for local density—estimated via kernel density estimation—and graph connectivity, enabling recursive generation of the clustering hierarchy. A key advantage of this method is its ability to automatically infer intrinsic thresholds, thereby substantially reducing reliance on manual parameter tuning. Extensive experiments on diverse heterogeneous benchmark datasets demonstrate that the proposed algorithm consistently outperforms state-of-the-art methods such as AChameleon and RNN-DBSCAN in both clustering accuracy and parameter robustness.
This work proposes an interactive, human-in-the-loop visual clustering framework for high-dimensional data, addressing the limitations of static dimensionality reduction methods that often lack interpretability and cannot incorporate human prior knowledge. By introducing a closed-loop feedback mechanism, the approach enables users to dynamically guide nonlinear projections through a small number of must-link and cannot-link constraints, while simultaneously refining low-dimensional embeddings via semi-supervised clustering. The framework further supports traceability from clustering results back to the original feature space, facilitating interpretable analysis. Experimental results on multiple benchmark datasets demonstrate that just a few rounds of user interaction significantly improve clustering quality, achieving both efficiency and interpretability in high-dimensional clustering tasks.
This study aims to identify clinically meaningful patient subgroups from electronic health records of breast cancer patients to uncover underlying pathological patterns. To this end, it introduces a novel approach that systematically integrates UMAP for nonlinear dimensionality reduction with DBSCAN for density-based clustering, applied across three independent real-world breast cancer datasets. Clustering quality is rigorously evaluated using a combination of internal validation metrics—DBCV, DCSI, and DISCO—to ensure robustness and interpretability. The proposed methodology substantially enhances the stability and clinical interpretability of the resulting clusters, successfully revealing multiple patient subgroups with distinct and significant clinical characteristics. These findings offer a data-driven foundation for advancing precision oncology and informing future mechanistic investigations into breast cancer heterogeneity.
This study addresses the lack of systematic evaluation regarding how dimensionality reduction methods influence clustering performance. Within a unified framework, it comprehensively assesses the impact of five dimensionality reduction techniques—PCA, Kernel PCA, VAE, Isomap, and MDS—across varying target dimensions on four mainstream clustering algorithms: k-means, Agglomerative Hierarchical Clustering (AHC), Gaussian Mixture Models (GMM), and OPTICS. Clustering quality is quantified using the Adjusted Rand Index (ARI). The work reveals, for the first time, the intricate coupling among intrinsic data geometry, dimensionality reduction strategy, and clustering efficacy, demonstrating that the choice of both reduction method and target dimension must be jointly tailored to the data’s underlying structure and the specific clustering algorithm. Indiscriminate application of dimensionality reduction can substantially degrade clustering performance.
This work addresses the limitations of traditional clustering algorithms—specifically, the high computational complexity (O(N²)) of density-based methods like DBSCAN and the inability of partitioning approaches such as K-Means to capture nonlinear structures or handle noise effectively. To overcome these challenges, the authors propose K-SCAN, a novel algorithm that synergistically integrates vector quantization with density-based analysis. K-SCAN first employs stochastic mini-batch K-Means to generate weighted micro-clusters and then performs density connectivity analysis on these micro-clusters. This approach achieves linear time complexity while accurately identifying nonlinear manifold structures and exhibiting robustness to noise. Experimental results demonstrate that K-SCAN runs over three times faster than BIRCH on million-scale datasets, attains an Adjusted Rand Index exceeding 0.99, and remains effective even with noise levels as high as 55%.