FastEnsemble: scalable ensemble clustering on large networks

📅 2024-09-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Community detection in large-scale networks suffers from result instability due to algorithmic randomness, heterogeneity across methods, sensitivity to resolution parameters, and partial coverage. Method: This paper proposes the first scalable ensemble clustering framework capable of handling ultra-large networks (>3 million nodes). It constructs a weighted consensus matrix from community partitions generated by multiple algorithms, across multiple resolution levels and random seeds, and applies fast spectral decomposition to derive robust consensus communities. Contribution/Results: The method achieves higher accuracy than ECG and FastConsensus on synthetic benchmarks; it is significantly faster than FastConsensus—enabling real-time analysis on networks with up to ten million nodes—and effectively mitigates the resolution limit problem inherent in modularity-based approaches.

Technology Category

Machine Learning: Ensemble MethodsData Mining & Knowledge Management: Graph Mining, Social Network Analysis & CommunitySearch and Optimization: Distributed Search

Application Category

Graph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsWeb Mining and Content Analysis: Robustness and generalizability of Web mining methodsResponsible Web: Human-perceived consequences of algorithmic deployment on the web
📝 Abstract
Many community detection algorithms are inherently stochastic, leading to variations in their output depending on input parameters and random seeds. This variability makes the results of a single run of these algorithms less reliable. Moreover, different clustering algorithms, optimization criteria (e.g., modularity, the Constant Potts model), and resolution values can result in substantially different partitions on the same network. Consensus clustering methods, such as ECG and FastConsensus, have been proposed to reduce the instability of non-deterministic algorithms and improve their accuracy by combining a set of partitions resulting from multiple runs of a clustering algorithm. In this work, we introduce FastEnsemble, a new consensus clustering method. Our results on a wide range of synthetic networks show that FastEnsemble produces more accurate clusterings than two other consensus clustering methods, ECG and FastConsensus, for many model conditions. Furthermore, FastEnsemble is fast enough to be used on networks with more than 3 million nodes, and so improves on the speed and scalability of FastConsensus. Finally, we showcase the utility of consensus clustering methods in mitigating the effect of resolution limit and clustering networks that are only partially covered by communities.
Problem

Research questions and friction points this paper is trying to address.

Improves reliability of community detection
Enhances accuracy of consensus clustering
Scales to networks with over 3 million nodes
Innovation

Methods, ideas, or system contributions that make the work stand out.

FastEnsemble consensus clustering method
Improves accuracy and scalability
Mitigates resolution limit effects
🔎 Similar Papers
No similar papers found.
University of Illinois Urbana-Champaign
Y
Yasamin Tabatabaee
Siebel School of Computing and Data Science, University of Illinois Urbana-Champaign, Urbana IL 61801
E
Eleanor Wedell
Siebel School of Computing and Data Science, University of Illinois Urbana-Champaign, Urbana IL 61801
Minhyuk Park
Minhyuk Park
Graduate Student, University of Illinois Urbana-Champaign
Tandy Warnow
Tandy Warnow
Grainger Distinguished Chair in Engineering, UIUC
Computer ScienceComputational BiologyPhylogeneticsMetagenomicsMultiple Sequence Alignment