🤖 AI Summary
This work addresses the problem of community detection in the two-community stochastic block model with multiple independent graph samples. It proposes averaging the adjacency matrices across samples followed by spectral partitioning to suppress noise and enhance recovery accuracy. Theoretically, the study establishes a spectral norm bound for the aggregated noise matrix under multiple samples and proves that the probability of community recovery error decays exponentially with the number of samples, providing the first rigorous guarantee for graph data augmentation in this setting. The analysis combines matrix averaging, a simplified spectral algorithm, and tools from the Davis–Kahan theorem and random matrix theory. Experiments on graphs with up to 1,000 nodes and as few as 2–3 samples demonstrate the tightness of the theoretical bounds and confirm that even a small number of samples significantly improves detection performance.
📝 Abstract
We study community detection in the two-block stochastic block model under the setting where multiple independent graph samples drawn from the same distribution are available. Building on a recently simplified spectral algorithm that preserves the independence of adjacency matrix entries throughout, we show that averaging $m$ independent samples before applying spectral partitioning reduces the error bound $γ$ exponentially in $m$: specifically, one can find a $γ$-correct partition with probability $1 - o(1)$ whenever $\frac{(a-b)^2}{a+b} \geq \frac{C}{m} \log \frac{2}γ$, improving the single-sample requirement by a factor of $m$. The key technical contribution is a multi-sample analogue of the spectral norm bound on the noise matrix, which propagates through the Davis-Kahan subspace angle analysis to yield the improved recovery guarantee. We provide experimental validation across a range of graph sizes ($n$ up to $1000$) and sample counts ($m$ up to $9$), demonstrating that the derived bounds are sharp and that even two or three samples yield dramatic improvements in recovery accuracy. Our results offer a rigorous theoretical foundation for graph data augmentation strategies used in modern graph representation learning.