🤖 AI Summary
Community detection in large-scale networks suffers from result instability due to algorithmic randomness, heterogeneity across methods, sensitivity to resolution parameters, and partial coverage. Method: This paper proposes the first scalable ensemble clustering framework capable of handling ultra-large networks (>3 million nodes). It constructs a weighted consensus matrix from community partitions generated by multiple algorithms, across multiple resolution levels and random seeds, and applies fast spectral decomposition to derive robust consensus communities. Contribution/Results: The method achieves higher accuracy than ECG and FastConsensus on synthetic benchmarks; it is significantly faster than FastConsensus—enabling real-time analysis on networks with up to ten million nodes—and effectively mitigates the resolution limit problem inherent in modularity-based approaches.
📝 Abstract
Many community detection algorithms are inherently stochastic, leading to variations in their output depending on input parameters and random seeds. This variability makes the results of a single run of these algorithms less reliable. Moreover, different clustering algorithms, optimization criteria (e.g., modularity, the Constant Potts model), and resolution values can result in substantially different partitions on the same network. Consensus clustering methods, such as ECG and FastConsensus, have been proposed to reduce the instability of non-deterministic algorithms and improve their accuracy by combining a set of partitions resulting from multiple runs of a clustering algorithm. In this work, we introduce FastEnsemble, a new consensus clustering method. Our results on a wide range of synthetic networks show that FastEnsemble produces more accurate clusterings than two other consensus clustering methods, ECG and FastConsensus, for many model conditions. Furthermore, FastEnsemble is fast enough to be used on networks with more than 3 million nodes, and so improves on the speed and scalability of FastConsensus. Finally, we showcase the utility of consensus clustering methods in mitigating the effect of resolution limit and clustering networks that are only partially covered by communities.