🤖 AI Summary
Existing clustering approaches suffer from paradigm fragmentation, ambiguous applicability, and limited interpretability when handling high-dimensional, dynamic, and heterogeneous data. Method: This study systematically unifies five major clustering paradigms—partitioning, hierarchical, data-stream, subspace, and network clustering—and establishes, for the first time, a cross-paradigm applicability mapping framework that rigorously defines their extension boundaries and adaptation conditions in dynamic data streams and heterogeneous networks. A comprehensive evaluation of representative algorithms (K-means, DBSCAN, AGNES, BIRCH, CLIQUE, and community detection methods) is conducted using silhouette coefficient, F1-score, and visualization analysis. Contribution/Results: We propose a reusable three-stage practical guideline—preprocessing, algorithm selection, and validation—and deliver an application atlas spanning 12 disciplines. The framework significantly improves clustering accuracy and interpretability on high-dimensional sparse data.
📝 Abstract
This paper provides a comprehensive exploration of data clustering, emphasizing its methodologies and applications across different fields. Traditional techniques, including partitional and hierarchical clustering, are discussed alongside other approaches such as data stream, subspace and network clustering, highlighting their role in addressing complex, high-dimensional datasets. The paper also reviews the foundational principles of clustering, introduces common tools and methods, and examines its diverse applications in data science. Finally, the discussion concludes with insights into future directions, underscoring the centrality of clustering in driving innovation and enabling data-driven decision making.