Score
Designs and implements pipelines that extract feature representations and mine recurring patterns from structural data, producing representations of individual structural objects suitable for unsupervised grouping. Builds and configures density-based hierarchical clustering using HDBSCAN to discover and evaluate unsupervised structural clusters, characterize cluster stability and noise points, and analyze cluster membership to interpret structural motifs.
Persistent homology-based clustering algorithms—particularly in topological data analysis (e.g., Mapper)—often suffer from sensitivity to manually tuned hyperparameters, limiting robustness and broad applicability. To address this, we propose AuToMATo, the first fully automated, persistent-homology-driven clustering algorithm requiring no user-specified parameters. It integrates ToMATo’s density-peak detection with bootstrap-based significance testing to automatically identify statistically significant modes in the density function, thereby eliminating all hyperparameter dependence. Grounded rigorously in persistent homology theory, AuToMATo ensures mathematical soundness while delivering cross-dataset robustness. Extensive experiments demonstrate that AuToMATo consistently outperforms state-of-the-art parameter-free methods and frequently surpasses optimally tuned parametric alternatives across multiple benchmarks. An open-source Python implementation—fully compatible with the scikit-learn API—is publicly available and already integrated into standard topological analysis workflows.
Analyzing and comparing non-hierarchical multi-scale clustering sequences remains challenging due to the lack of stable, topology-aware representations. Method: We propose Multi-scale Clustering Filtration (MCF), a stable simplicial filtration that encodes clustering partitions at arbitrary scales. We systematically introduce persistent homology to this task by constructing MCF and its equivalent nerve complex, proving that in the hierarchical case it reduces to the Vietoris–Rips filtration on an ultrametric space. Contribution/Results: Empirical validation on synthetic data demonstrates that zero- and higher-dimensional persistence diagrams derived from MCF serve as robust topological features, effectively characterizing and distinguishing diverse multi-scale clustering structures. MCF thus establishes a novel paradigm for the quantitative evaluation and comparative analysis of multi-scale clustering, enabling principled, topology-driven assessment beyond traditional metrics.
Traditional density-based clustering algorithms (e.g., DBSCAN, OPTICS) suffer from prohibitive computational cost and memory infeasibility on high-dimensional, large-scale datasets. Method: This paper proposes sDBSCAN and sOPTICS—novel scalable variants that integrate theoretically grounded random projections to preserve neighborhood structure under cosine and other distance metrics; further introduces neighborhood-aware pruning and indexing optimizations to efficiently approximate core points and their neighborhoods within modified DBSCAN/OPTICS frameworks. Contribution/Results: The algorithms guarantee provably convergent approximate clustering structures and support interactive hierarchical exploration. Evaluated on million-scale real-world datasets, they achieve speedups of over an order of magnitude versus scikit-learn’s implementations (finishing in minutes), maintain controlled memory footprint, surpass state-of-the-art methods in clustering accuracy, and scale to ultra-large datasets beyond the reach of conventional density-based approaches.
Traditional correlation clustering struggles to model hierarchical structures and handle heterogeneous pairwise similarities—both positive and negative—within a unified framework. Method: We propose Hierarchical Correlation Clustering (HCC), the first correlation clustering formulation explicitly designed for hierarchical settings. HCC jointly optimizes cluster hierarchy and low-dimensional representations via an unsupervised tree-preserving embedding scheme. Furthermore, we generalize the minimax distance metric to correlation clustering, yielding a robust, noise-resilient dissimilarity measure tailored to heterogeneous pairwise constraints. Contribution/Results: Extensive experiments on standard benchmarks demonstrate that HCC significantly improves hierarchical clustering quality, embedding fidelity, and downstream task performance over state-of-the-art methods. By unifying hierarchy learning, representation learning, and robust constraint modeling in an unsupervised setting, HCC establishes a novel paradigm for unsupervised hierarchical representation learning.
Multi-parameter hierarchical clustering suffers from poor stability under data perturbations and lacks principled methods to extract stable single-parameter representations. Method: This paper proposes Persistable—a framework grounded in multiparameter persistent homology and the degree-Rips bifiltration—that yields density-aware, multiscale-consistent hierarchical clustering. It introduces the correspondence-interleaving distance to rigorously quantify hierarchical clustering stability, proves that the multiparameter degree-Rips bifiltration is stable, and demonstrates the instability of conventional single-parameter slices. The framework further proposes a novel density-adaptive stable slicing method and a theoretically guaranteed stable flattening algorithm. Results: On benchmark datasets, Persistable accurately recovers multiscale cluster structures. All core components—slicing, flattening, and parameter selection—are supported by formal stability and consistency guarantees, establishing the first theoretically robust pipeline for stable multiparameter hierarchical clustering.
This work addresses the limitations of traditional hierarchical clustering, which relies solely on pairwise distances and struggles to capture density variations and local connectivity inherent in graph-structured data. The authors propose a novel integrated approach that synergistically combines hierarchical, density-based, and graph clustering paradigms. Specifically, they construct KNN-induced subgraphs and introduce a cluster similarity measure that jointly accounts for local density—estimated via kernel density estimation—and graph connectivity, enabling recursive generation of the clustering hierarchy. A key advantage of this method is its ability to automatically infer intrinsic thresholds, thereby substantially reducing reliance on manual parameter tuning. Extensive experiments on diverse heterogeneous benchmark datasets demonstrate that the proposed algorithm consistently outperforms state-of-the-art methods such as AChameleon and RNN-DBSCAN in both clustering accuracy and parameter robustness.
This work proposes an interactive, human-in-the-loop visual clustering framework for high-dimensional data, addressing the limitations of static dimensionality reduction methods that often lack interpretability and cannot incorporate human prior knowledge. By introducing a closed-loop feedback mechanism, the approach enables users to dynamically guide nonlinear projections through a small number of must-link and cannot-link constraints, while simultaneously refining low-dimensional embeddings via semi-supervised clustering. The framework further supports traceability from clustering results back to the original feature space, facilitating interpretable analysis. Experimental results on multiple benchmark datasets demonstrate that just a few rounds of user interaction significantly improve clustering quality, achieving both efficiency and interpretability in high-dimensional clustering tasks.
This work proposes a hierarchical clustering algorithm grounded in topological data analysis to address the challenge of identifying clusters of arbitrary shape and salient outliers without requiring prior assumptions about data distribution. The method operates under any distance metric and eschews distributional assumptions by constructing a topological hierarchy and incorporating persistence analysis to rigorously assess cluster stability and outlier significance. Experimental evaluations on real-world datasets from domains such as image analysis, healthcare, and economics demonstrate that the algorithm yields semantically coherent and robust clustering results even in complex scenarios where conventional approaches fail, substantially enhancing the detection of non-convex structures and critical anomalous points.
This work addresses the limitations of traditional clustering algorithms—specifically, the high computational complexity (O(N²)) of density-based methods like DBSCAN and the inability of partitioning approaches such as K-Means to capture nonlinear structures or handle noise effectively. To overcome these challenges, the authors propose K-SCAN, a novel algorithm that synergistically integrates vector quantization with density-based analysis. K-SCAN first employs stochastic mini-batch K-Means to generate weighted micro-clusters and then performs density connectivity analysis on these micro-clusters. This approach achieves linear time complexity while accurately identifying nonlinear manifold structures and exhibiting robustness to noise. Experimental results demonstrate that K-SCAN runs over three times faster than BIRCH on million-scale datasets, attains an Adjusted Rand Index exceeding 0.99, and remains effective even with noise levels as high as 55%.
This work addresses the limitations of traditional batch clustering algorithms—such as HDBSCAN—in meeting the demands of real-time, memory-efficient, and dynamically adaptive narrative monitoring on social media. To overcome these challenges, we propose a three-stage pipeline that employs online density-based clustering for multilingual document streams. We further introduce a comprehensive evaluation framework that integrates conventional clustering metrics (e.g., Silhouette coefficient, Davies–Bouldin index) with narrative-specific indicators, including narrative distinctiveness, contingency, and variance. Through sliding-window simulations, we systematically compare multiple incremental clustering algorithms and identify an optimal solution that balances clustering quality, computational efficiency, and memory footprint. The resulting approach effectively supports real-time tracking of narrative evolution and demonstrates strong practical viability for deployment in operational settings.