🤖 AI Summary
Existing spectral estimators for random dot product graphs (RDPGs) ignore sampling likelihood information, resulting in statistical inefficiency.
Method: We propose a separable, log-concave surrogate likelihood function that accurately approximates the exact likelihood while preserving computational tractability and restoring likelihood-based inference.
Contribution/Results: This work provides the first rigorously justified surrogate likelihood for RDPGs. Under the frequentist framework, we establish existence, uniqueness, and asymptotic normality of the resulting estimator. Under the Bayesian framework, we prove a Bernstein–von Mises theorem, ensuring that posterior credible sets achieve valid frequentist coverage. Empirically, our method achieves significantly lower estimation error than spectral methods—demonstrated on both synthetic benchmarks and real-world Wikipedia network data. An open-source R package, *lgraph*, implements the methodology for immediate practical use.
📝 Abstract
Spectral estimators have been broadly applied to statistical network analysis but they do not incorporate the likelihood information of the network sampling model. This paper proposes a novel surrogate likelihood function for statistical inference of a class of popular network models referred to as random dot product graphs. In contrast to the structurally complicated exact likelihood function, the surrogate likelihood function has a separable structure and is log-concave yet approximates the exact likelihood function well. From the frequentist perspective, we study the maximum surrogate likelihood estimator and establish the accompanying theory. We show its existence, uniqueness, large sample properties, and that it improves upon the baseline spectral estimator with a smaller sum of squared errors. A computationally convenient stochastic gradient descent algorithm is designed for finding the maximum surrogate likelihood estimator in practice. From the Bayesian perspective, we establish the Bernstein– von Mises theorem of the posterior distribution with the surrogate likelihood function and show that the resulting credible sets have the correct frequentist coverage. The empirical performance of the proposed surrogate-likelihood-based methods is validated through the analyses of simulation examples and a realworld Wikipedia graph dataset. An R package implementing the proposed computation algorithms is publicly available at https://fangzheng-xie.github.io./materials/lgraph_0.1.0.tar.gz.