🤖 AI Summary
This work investigates the information-theoretic limits and efficient algorithms for community recovery in multilayer stochastic block models under constant average degree and high-dimensional node covariates, where the covariate dimension scales proportionally with the number of nodes. By introducing a novel Bernoulli–Gaussian moment comparison inequality, the authors establish a statistical version of the “reduction-to-chi-square-divergence” framework. Combined with decorated cycle and path counting algorithms, this approach rigorously characterizes the phase transition threshold for weak recovery. The theoretical analysis demonstrates that the proposed algorithm achieves information-theoretic optimality at this threshold, with no statistical–computational gap. This result constitutes the first proof of equivalence between statistical and computational limits in this high-dimensional, sparse, multilayer setting.
📝 Abstract
We consider the problem of community detection from the joint observation of a high-dimensional covariate matrix and $L$ sparse networks, all encoding noisy, partial information about the latent community labels of $n$ subjects. In the asymptotic regime where the networks have constant average degree and the number of features $p$ grows proportionally with $n$, we derive a sharp threshold under which detecting and estimating the subject labels is possible. Our results extend the work of \cite{MN23} to the constant-degree regime with noisy measurements, and also resolve a conjecture in \cite{YLS24+} when the number of networks is a constant. Our information-theoretic lower bound is obtained via a novel comparison inequality between Bernoulli and Gaussian moments, as well as a statistical variant of the ``recovery to chi-square divergence reduction''argument inspired by \cite{DHSS25}. On the algorithmic side, we design efficient algorithms based on counting decorated cycles and decorated paths and prove that they achieve the sharp threshold for both detection and weak recovery. In particular, our results show that there is no statistical-computational gap in this setting.