Score
Designs, implements, and evaluates algorithms and software that identify groups of nodes (communities) in graph-structured data, producing partitions or overlapping/hierarchical clusterings of vertices. Builds and analyzes methods for choosing objective functions and similarity measures, scaling and optimizing performance, and measuring community quality and stability using appropriate metrics and benchmarks.
Core-periphery structure lacks a unified definition and standardized detection methodology, leading to conceptual ambiguity and inconsistent evaluation—hindering both theoretical advancement and practical application. This paper addresses this gap through a systematic literature review and methodological comparison, integrating graph-theoretic modeling, clustering algorithms, and structural evaluation metrics to classify, empirically benchmark, and delineate the boundaries of mainstream core-periphery detection methods. It clarifies their distinctions from and relationships with community structure, along with contextual applicability conditions. The study establishes a comprehensive theoretical framework encompassing conceptual foundations, a taxonomy of methods, and principled evaluation criteria; identifies key open challenges; and proposes a standardized definition and a reproducible, metric-driven assessment protocol. These contributions provide a systematic foundation for algorithm design, cross-method comparison, and empirical analysis of real-world networks.
In practical applications, community detection methods lack standardized evaluation protocols, and their impact on downstream graph mining tasks is often overlooked. This paper systematically investigates how diverse community detection algorithms affect the performance of link prediction and node classification. We propose a unified, extensible evaluation framework that integrates structured community feature extraction, statistical analysis, and machine learning modeling to enable cross-algorithm performance comparison. Experimental results across multiple benchmark datasets demonstrate that algorithm selection significantly influences downstream task accuracy, with distinct methods exhibiting pronounced strengths and weaknesses depending on the specific task. Our framework provides reproducible, empirically grounded guidance for selecting appropriate community detection methods tailored to concrete application scenarios, thereby bridging the gap between community detection research and real-world graph analytics. (149 words)
Existing stochastic block model (SBM)-based community detection methods frequently yield internally disconnected or weakly connected communities, compromising structural fidelity and clustering accuracy. Method: We propose Well-Connected Clusters (WCC), the first approach to systematically uncover and address graph-tool’s implicit preference for disconnected partitions—arising from its description length formulation—via edge-cut analysis and explicit connectivity constraints integrated into the optimization objective. WCC is compatible with both flat and nested SBM variants. Contribution/Results: Evaluated on synthetic benchmarks and large-scale real-world networks (up to million-node scale), WCC reveals pervasive disconnection issues across all baseline SBM implementations; graph-tool outperforms PySBM, yet still suffers from poor connectivity. WCC significantly improves modularity, normalized mutual information (NMI), and connectivity metrics while preserving model parsimony—establishing the first general-purpose, connectivity-guaranteed enhancement for the SBM framework.
Existing dynamic network benchmarks largely neglect community evolution modeling, thus failing to adequately evaluate algorithms’ capability to capture realistic community lifecycles (e.g., growth, contraction, splitting, merging, dissolution) and node-level dynamics (e.g., emergence, disappearance, inter-community migration). Method: We propose a community-centric temporal network generation model that—uniquely—couples community evolution with node dynamics, enabling fine-grained specification of ground-truth community trajectories. Leveraging stochastic graph mechanisms and lifecycle-aware control policies, we construct a scalable benchmark suite, accompanied by standardized evaluation metrics and interactive visualization tools. Contribution/Results: Extensive experiments demonstrate the benchmark’s sensitivity to three state-of-the-art dynamic community detection algorithms, significantly enhancing quantitative assessment of evolutionary behavior recognition and member trajectory tracking. Our work fills a critical gap in dynamic community tracking evaluation, providing the first principled, configurable, and empirically validated benchmark for this task.
Traditional divisive algorithms for community detection in complex networks suffer from sensitivity to initial edge betweenness and a tendency to converge prematurely to local optima of the modularity metric (Q). Method: This paper proposes a (Q)-driven improved divisive algorithm that deeply integrates modularity optimization throughout the entire splitting process. Key components include dynamic edge-weight pruning, adaptive threshold adjustment for partitioning, enhanced edge-betweenness computation, incremental (Q) evaluation, iterative split-backtrack optimization, and a (Q)-guided termination mechanism. Contribution/Results: The method significantly improves partitioning accuracy and robustness. On standard benchmark networks, it achieves an average modularity gain of 1.2%–3.7% over state-of-the-art approaches—including Girvan–Newman (GN) and Fast Newman—yielding communities with stronger internal cohesion and weaker inter-community coupling.
To address the challenge of modeling higher-order, dynamic, and overlapping multi-body interactions in temporal hypergraphs, this paper proposes an unsupervised pattern discovery method based on hyperedge clustering. Unlike conventional node-centric clustering, our approach defines a time-aware structural similarity measure in the hyperedge space—featuring three scalable design variants—and integrates spectral clustering with hyperedge-space embedding to automatically identify dense interaction subpatterns. This work is the first to extend the edge-clustering paradigm to temporal hypergraphs, overcoming inherent limitations of node-based modeling in representing higher-order structural dependencies. Experiments on large-scale collaborative hypergraphs demonstrate that the discovered patterns exhibit strong semantic coherence and interpretability, and effectively support downstream tasks such as collaboration prediction and role discovery.
Existing hypergraph partitioning methods often become trapped in local optima, limiting partition quality. This work proposes ComPart, a novel framework that integrates community structure guidance during both the initial partitioning and uncoarsening phases. It is the first to comprehensively incorporate community detection throughout the entire uncoarsening process and extends the theory of local dense decomposition from graphs to hypergraphs to generate high-quality initial partitions. By synergistically combining diverse community detection techniques, hypergraph local dense decomposition, and a multilevel partitioning strategy, ComPart consistently outperforms state-of-the-art methods on standard benchmarks, achieving significantly improved partition quality.
This work addresses the high computational cost of community detection in dynamic graphs by proposing an efficient GPU-accelerated method for temporal network analysis. Leveraging the NVIDIA RAPIDS ecosystem, it presents the first GPU-based implementation of modularity optimization and symmetric Bethe–Hessian operator eigendecomposition across multiple graph snapshots. The approach integrates the Leiden algorithm with Dask’s distributed scheduling framework to enable scalable community tracking in large-scale temporal networks. Designed with compatibility for the NetworkX-Temporal interface, the system seamlessly fits into existing graph analytics pipelines. Experimental results demonstrate up to a 1000× speedup over CPU baselines under equivalent workloads, substantially enhancing the efficiency of temporal network analysis in domains such as epidemic spreading, financial systems, and cybersecurity.
Existing methods for comparing graph partitions often disregard the underlying graph topology, making it difficult to accurately capture the cohesion within and separation between communities. This work proposes a graph-aware distance framework that constructs a topology-respecting metric by inducing edge partitions, enabling meaningful comparison of continuous graph partitions. The framework adheres to a local graph-aware refinement criterion and is theoretically shown—under the stochastic block model—to be sensitive to topological perturbations. Specifically, the proposed distance almost surely increases with stronger perturbations in both intra- and inter-community splitting scenarios, significantly outperforming conventional metrics such as variation of information, van Dongen distance, and binary cut distance. This provides a structurally consistent and topology-sensitive criterion for evaluating graph partitions.
This work investigates the asymptotic behavior of modularity $q^*(G)$ for Erdős–Rényi random graphs $G_{n,p}$, aiming to characterize the fundamental limits of modularity as a community structure metric. Using probabilistic graph theory and extremal analysis, we establish almost sure convergence of $q^*(G_{n,p})$ across distinct regimes of edge probability $p$, and derive the first tight asymptotic bounds: when $p = o(1)$ and $np o infty$, $q^*(G_{n,p}) = 1 - o(1)$; when $p$ is constant, $q^*(G_{n,p}) = Theta(1/sqrt{n})$. These results demonstrate that modularity effectively detects community structure in sparse graphs but deteriorates in dense regimes. The analysis reveals an intrinsic resolution limit of modularity—its inability to resolve fine-grained communities as density increases—and provides a rigorous theoretical benchmark for evaluating and designing community detection algorithms, including a provable performance ceiling.