Score
Design, implement, and evaluate algorithms and systems that identify cohesive groups or clusters in relational and feature data—principally graph-structured networks and feature/text corpora—using methods such as graph- and spectral-clustering, greedy and hierarchical algorithms, density-based (HDBSCAN), model- and similarity-based, semantic/thematic, and adaptive deep clustering. This work includes building multi-layer network representations, implementing and scaling clustering/search/pruning procedures to find near‑optimal node- or user-centric partitions, and integrating clustering with segmentation and downstream analysis.
Core-periphery structure lacks a unified definition and standardized detection methodology, leading to conceptual ambiguity and inconsistent evaluation—hindering both theoretical advancement and practical application. This paper addresses this gap through a systematic literature review and methodological comparison, integrating graph-theoretic modeling, clustering algorithms, and structural evaluation metrics to classify, empirically benchmark, and delineate the boundaries of mainstream core-periphery detection methods. It clarifies their distinctions from and relationships with community structure, along with contextual applicability conditions. The study establishes a comprehensive theoretical framework encompassing conceptual foundations, a taxonomy of methods, and principled evaluation criteria; identifies key open challenges; and proposes a standardized definition and a reproducible, metric-driven assessment protocol. These contributions provide a systematic foundation for algorithm design, cross-method comparison, and empirical analysis of real-world networks.
This work addresses the challenge of effectively integrating graph structure and node attributes for unsupervised clustering in attributed graphs. The authors propose a multi-round self-training framework based on graph neural networks that alternately refines node representations and cluster assignments in an unsupervised manner. In each round, the current clustering result is used to reconstruct the graph structure, which is then fused with the original graph to form a context-aware graph for generating improved representations. The key innovation lies in the dynamic, synergistic utilization of both edge structure and node attributes, overcoming limitations of conventional single-round training or reliance on a single information source. Experiments demonstrate that the method significantly outperforms baselines using only structure or attributes on synthetic data, that multi-round learning surpasses extended single-round training, and that it achieves state-of-the-art performance on real-world datasets under balanced clustering scenarios.
This work addresses the limitations of traditional hierarchical clustering, which relies solely on pairwise distances and struggles to capture density variations and local connectivity inherent in graph-structured data. The authors propose a novel integrated approach that synergistically combines hierarchical, density-based, and graph clustering paradigms. Specifically, they construct KNN-induced subgraphs and introduce a cluster similarity measure that jointly accounts for local density—estimated via kernel density estimation—and graph connectivity, enabling recursive generation of the clustering hierarchy. A key advantage of this method is its ability to automatically infer intrinsic thresholds, thereby substantially reducing reliance on manual parameter tuning. Extensive experiments on diverse heterogeneous benchmark datasets demonstrate that the proposed algorithm consistently outperforms state-of-the-art methods such as AChameleon and RNN-DBSCAN in both clustering accuracy and parameter robustness.
Traditional methods for node classification and clustering on graph-structured data—such as social and biological networks—are limited by their inability to effectively model non-Euclidean geometric properties inherent in graphs. To address this, we propose a synergistic modeling framework that integrates classical graph algorithms with graph neural networks (GNNs). Our approach systematically analyzes the representational disparities between these two paradigms and constructs an interpretable graph representation learning scheme, thereby providing theoretical foundations for non-Euclidean structural modeling. Extensive experiments on multiple benchmark graph datasets demonstrate that the proposed method achieves 43%–70% higher accuracy in both node classification and clustering compared to standalone classical algorithms (e.g., Label Propagation, Spectral Clustering) and baseline GNN models. These results robustly validate the dual advantages of our fusion strategy—superior accuracy and enhanced robustness—over existing approaches.
Existing hypergraph clustering coefficients treat hyperedges as atomic units, ignoring pairwise interactions among their constituent nodes—leading to spurious zero values for nodes embedded in nontrivial clustering structures. Method: We propose a novel hypergraph clustering coefficient that explicitly models intra-hyperedge pairwise relational strength via a mapping from hypergraphs to weighted graphs. Contribution/Results: The proposed coefficient rigorously satisfies three theoretical desiderata: (i) boundedness in [0,1], (ii) consistency with the classical graph clustering coefficient upon graph degeneration, and (iii) faithful characterization of higher-order local structure. Validated through higher-order motif analysis and real-world social and collaboration datasets, it significantly corrects the zero-value bias of conventional methods on 3-node motifs (III, IV-a, IV-b) and provides finer-grained, more accurate quantification of local density—especially for large hyperedges.
This work addresses the challenge of simultaneously performing network-level clustering and node-level community detection in multi-network analysis. We propose the first Bayesian nonparametric model based on the nested Dirichlet process (NDP), enabling joint inference of both the number of network types and the number of communities within each network. The model accommodates unlabeled, structurally heterogeneous networks with unequal node sets—overcoming key limitations in modeling anonymized nodes and scale-heterogeneous networks. We develop three Gibbs samplers—standard, collapsed, and blocked—to ensure efficient posterior inference. Extensive experiments demonstrate that the method accurately recovers hierarchical clustering structures on synthetic data and achieves superior performance on two real-world social network datasets. To our knowledge, this is the first unified, adaptive, and scalable framework for multi-network co-analysis, offering principled uncertainty quantification and automatic complexity control without requiring prespecified numbers of clusters or communities.
In practical applications, community detection methods lack standardized evaluation protocols, and their impact on downstream graph mining tasks is often overlooked. This paper systematically investigates how diverse community detection algorithms affect the performance of link prediction and node classification. We propose a unified, extensible evaluation framework that integrates structured community feature extraction, statistical analysis, and machine learning modeling to enable cross-algorithm performance comparison. Experimental results across multiple benchmark datasets demonstrate that algorithm selection significantly influences downstream task accuracy, with distinct methods exhibiting pronounced strengths and weaknesses depending on the specific task. Our framework provides reproducible, empirically grounded guidance for selecting appropriate community detection methods tailored to concrete application scenarios, thereby bridging the gap between community detection research and real-world graph analytics. (149 words)
This work addresses the challenge of simultaneously identifying community structure and recovering its underlying hierarchical organization in networks, a task for which existing methods often rely on pre-specified parameters. The authors propose NHC-TST, a top-down, fully data-driven hierarchical clustering algorithm that formalizes hierarchy via a hierarchical distance matrix and recursively partitions the network using spectral clustering. An adaptive stopping criterion based on the graph-based two-sample test eliminates the need to predefine the number of clusters or tree depth, enabling reconstruction of unbalanced hierarchical trees. Theoretical analysis establishes the algorithm’s statistical consistency and structural recovery accuracy. Experiments demonstrate that NHC-TST precisely recovers both cluster memberships and hierarchical relationships across diverse synthetic networks and uncovers multi-scale dynamic structures in global migration data that flat clustering approaches fail to capture.
Spectral clustering lacks rigorous theoretical analysis for graphs with hierarchical or directed structures. Method: We propose a general performance criterion based on spectral gaps: accurate recovery of multiscale and directionally coherent clusters is guaranteed when the smallest eigenvalues of a Hermitian matrix representation form well-separated groups from the rest of the spectrum. Our approach transcends traditional Laplacian-based frameworks by extending spectral clustering theory to arbitrary Hermitian matrix representations—including Hermitianized formulations for directed graphs—and integrates symmetric graph representations with spectral graph theory via eigenvalue decomposition and spectral gap analysis. Contribution/Results: The resulting theory yields verifiable, interpretable guarantees. Experiments demonstrate that it accurately predicts clustering performance on synthetic benchmarks and real-world ecological networks (e.g., trophic level inference), significantly enhancing the interpretability and applicability of spectral clustering on complex-structured graphs.
This study addresses the absence of a unified theoretical framework for network partitioning clustering that accounts for both hard/soft assignment mechanisms and non-vertex-centered cluster prototypes. The authors systematically analyze four classical models: the hard-assignment p-median problem (PMP) and spectral clustering (SSC), alongside the soft-assignment probabilistic density clustering (PDC) and fuzzy c-means (FCM). They reveal, for the first time, that optimal solutions of PMP and PDC are inherently confined to graph vertices, whereas SSC and FCM can yield cluster centers located along edges. Through rigorous mathematical analysis of their structural properties and optimization behaviors under network topology, the work elucidates the critical roles of bottleneck points and vertex-constrained solutions, thereby establishing a theoretical foundation for efficient clustering algorithms in facility location, network design, and similarity search via graph embeddings.
Density-based clustering often suffers from sensitivity to manually tuned hyperparameters—such as density thresholds or minimum cluster size—especially when prior knowledge about data distribution is unavailable, leading to poor robustness. To address this, we propose PLSCAN, the first algorithm integrating scale-space clustering with persistent homology theory to construct an adaptive metric space that automatically identifies stable leaf clusters of HDBSCAN* across all scales—without any hyperparameter tuning. Our method leverages hierarchical density estimation via mutual reachability distance, scale-space trajectory tracking, and extraction of persistent leaf nodes. Experiments demonstrate that PLSCAN achieves higher average Adjusted Rand Index (ARI) than HDBSCAN* on multiple real-world datasets and exhibits superior robustness to perturbations in neighborhood size. Its computational efficiency is comparable to k-Means in low dimensions and matches HDBSCAN* in high dimensions. The core contribution is a fully parameter-free framework for extracting multi-scale density structures.