Two-Sample Testing for Inhomogeneous Random Graphs in Non-Integral $L_r$ Norms

πŸ“… 2026-09-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the theoretical gap of mismatched upper and lower bounds in two-sample testing for heterogeneous random graphs under non-integer $L_r$ norms. We propose a combined-statistic testing method based on HΓΆlder interpolation, which exploits the complementary properties of second- and higher-order statistics to construct adaptive thresholds for joint hypothesis testing, thereby unifying statistical thresholds across different orders. Theoretically, we prove that this test achieves the conjectured rate, establishing minimax optimal sample complexity bounds for any fixed $r \ge 1$. Empirically, simulations validate the theoretically predicted convergence rates and reveal the advantageous regimes of individual statistics under varying sparsity levels.
πŸ“ Abstract
Testing whether two populations of networks share the same edge probabilities is a basic problem in network inference. How hard it is depends on the norm used to measure the difference. For the inhomogeneous Erd\H{o}s--R\'enyi (IER) model, the optimal sample complexity is known for every integer $L_r$ norm and for $1\le r<2$. For non-integral $r>2$, however, the known upper and lower bounds do not match, and the lower bound was conjectured to be tight. We study this gap for two-sample testing on aligned vertices. We propose a test that runs two published statistics, of orders $2$ and $\lceil r\rceil$, on the same data and rejects if either one rejects. Its thresholds come from H\"older interpolation, so that both statistics have the same sample cost. We prove that this test attains the conjectured rate. Combined with earlier results, this shows that for every fixed $r\ge1$ the minimax sample complexity is of order $n^{\max\{4/r-1,\,2/r\}}/\epsilon^2$, even when the separation changes with $n$. In simulations with $n$ between 32 and 256, the number of graphs needed for 80\% power at level $0.05$ grows with $n$ at a rate consistent with the theory. For $r=2.5$, for example, the fitted exponent is $0.78$, against the theoretical value $0.8$. Interestingly, the two statistics split the work as the interpolation argument suggests: the higher-order statistic is more powerful when only a few edges change, and the $L_2$ statistic when many edges change.
Problem

Research questions and friction points this paper is trying to address.

Two-sample testing
Inhomogeneous random graphs
Non-integral L_r norms
Sample complexity
Network inference
πŸ”Ž Similar Papers