Score
Designs and analyzes estimators that use counts of cycles in graphs—typically forming log‑ratios of cycle counts—to estimate ratio‑type parameters such as node‑level β or other relative quantities. Builds explicit closed‑form CCR estimators adapted to very sparse networks and evaluates their statistical properties, e.g., mean squared error and minimax optimality.
This study addresses the challenge of estimating node-specific parameters in the $\beta$-model for sparse networks. The authors propose a Cycle Counting Ratio (CCR) estimator based on log-ratios of cycle counts, which leverages subgraph (cycle) enumeration to construct an explicitly computable statistic. Under extremely weak sparsity conditions—down to network density of order $\log n / n$—the method establishes, for the first time, consistency, uniform consistency, asymptotic normality, and minimax optimal convergence rates for individual node parameter estimation. Theoretical analysis demonstrates that the CCR estimator achieves the optimal rate in mean squared error, while numerical experiments and real-world sparse network data confirm its superior empirical performance.
Existing spectral estimators for random dot product graphs (RDPGs) ignore sampling likelihood information, resulting in statistical inefficiency. Method: We propose a separable, log-concave surrogate likelihood function that accurately approximates the exact likelihood while preserving computational tractability and restoring likelihood-based inference. Contribution/Results: This work provides the first rigorously justified surrogate likelihood for RDPGs. Under the frequentist framework, we establish existence, uniqueness, and asymptotic normality of the resulting estimator. Under the Bayesian framework, we prove a Bernstein–von Mises theorem, ensuring that posterior credible sets achieve valid frequentist coverage. Empirically, our method achieves significantly lower estimation error than spectral methods—demonstrated on both synthetic benchmarks and real-world Wikipedia network data. An open-source R package, *lgraph*, implements the methodology for immediate practical use.
This paper addresses the challenging problem of automatically estimating the number of clusters $k$ in high-dimensional data. We propose a nonparametric method based on similarity graphs: it constructs a robust graph structure, extracts intrinsic similarity patterns via spectral analysis of the graph Laplacian, and introduces an analytically tractable graph statistic—marking the first approach to directly leverage graph-structural features for $k$ selection. We establish theoretical consistency of the estimator under high-dimensional asymptotics. Unlike existing methods, our framework imposes no assumptions on underlying data distributions or specific clustering algorithms, and is both dimensionality-agnostic and computationally efficient. Extensive experiments on synthetic benchmarks, medical imaging, and RNA-seq datasets demonstrate substantial improvements in estimation accuracy under high-dimensional settings, along with enhanced robustness and generalizability.
This work investigates the information-theoretic limits and efficient algorithms for community recovery in multilayer stochastic block models under constant average degree and high-dimensional node covariates, where the covariate dimension scales proportionally with the number of nodes. By introducing a novel Bernoulli–Gaussian moment comparison inequality, the authors establish a statistical version of the “reduction-to-chi-square-divergence” framework. Combined with decorated cycle and path counting algorithms, this approach rigorously characterizes the phase transition threshold for weak recovery. The theoretical analysis demonstrates that the proposed algorithm achieves information-theoretic optimality at this threshold, with no statistical–computational gap. This result constitutes the first proof of equivalence between statistical and computational limits in this high-dimensional, sparse, multilayer setting.
Scalable random models for high-dimensional topological data analysis (TDA) remain scarce, particularly for abstract cell complexes (CCs). Method: We introduce the first Erdős–Rényi–style random CC model, constructing complexes layerwise by dimension via probabilistic cell addition, with a focus on overcoming sampling bottlenecks in the 2D case. Our approach features a novel boundary-length-constrained 2-cell sampling mechanism and a fast, enumeration-free estimator for the number of simple cycles in a graph. By integrating probabilistic graphical modeling, combinatorial approximation, and randomized algorithms, we achieve tunable-distribution sampling of 2D random CCs. Contribution/Results: We release py-raccoon, an open-source toolkit implementing the model. Experiments validate its effectiveness as a null model in TDA and as a graph augmentation tool for topological feature enhancement, demonstrating both theoretical soundness and practical utility.
This work proposes a statistical inference framework based on a hidden Markov network model to accurately estimate subgraph densities and enable joint multi-timepoint comparisons for non-i.i.d. dynamic network sequences subject to observation errors. By explicitly modeling edge-wise observation noise, the method achieves, for the first time, robust inference of subgraph densities across heterogeneous network snapshots and leverages information from multiple time points to enhance estimation efficiency. Theoretical analysis demonstrates that the proposed approach enjoys favorable asymptotic properties in large-scale networks, substantially improving both accuracy and computational efficiency in inferring subgraph structures from noisy dynamic networks.
This work addresses the problem of efficiently approximating the count of 4-cycles in graph edge streams presented in arbitrary order. To this end, the authors propose two induced subgraph sampling–based algorithms: a two-pass algorithm that achieves theoretically optimal space complexity on graphs with bounded degeneracy, marking the first optimal solution in this streaming model; and a single-pass algorithm tailored for sparse networks where 4-cycles are uniformly distributed, demonstrating robust performance on non-bipartite graphs such as social networks. Experimental evaluation shows that the two-pass algorithm significantly outperforms existing methods on real-world graph streams, while the single-pass variant also exhibits strong performance in its intended application scenarios.
In high-dimensional network analysis, when the number of observations is much smaller than the number of variables, model misspecification often leads to an inflated false positive rate in neighborhood estimation. This work proposes a conservative neighborhood selection method based on penalizing the volume of the model space, extending the Minimum Description Length (MDL) principle to high-dimensional settings and integrating ridge regression to elucidate its impact on mean squared error. The approach is applicable under both linear and nonlinear true models and theoretically guarantees either consistent neighborhood recovery under correct specification or a sparser—yet safer—estimate under misspecification, thereby substantially reducing false positive edges. This strategy overcomes key limitations of conventional methods such as Lasso and information criteria like AIC and BIC.
This work investigates the problem of distinguishing between two structured generative models under the planted-vs-planted setting with vanishing error probability, focusing on community-counting tasks in the planted submatrix and dense subgraph models. We propose a unified analytical framework based on low-degree polynomial tests, introducing a latent-variable expansion technique inspired by low-degree recovery and incorporating signal–noise separation and pruning strategies. For the first time, we establish sharp thresholds for low-degree testing in such problems: strong detection exhibits a sharp phase transition, whereas weak detection displays a smooth transition. Our derived upper and lower bounds match precisely, and the resulting detection threshold aligns—up to constant factors—with the known low-degree recovery threshold.
This work addresses the tendency of Laplacian-constrained graphical models—such as the Laplacian-constrained Gaussian graphical model (LCGGM) and the Hüsler–Reiss model—to yield overly dense graphs during structure learning, which compromises interpretability and scalability. The paper introduces, for the first time, spectral graph sparsification as a post-processing step: without requiring additional hyperparameter tuning, it replaces the original Laplacian estimate with a spectrally approximated sparse Laplacian and refits the model. This approach effectively enhances both the sparsity and accuracy of the estimated graph structure. Integrating spectral graph theory, Laplacian-constrained Gaussian graphical models, extreme-value graphical models, and graph sparsification techniques, the method demonstrates superior performance on Erdős–Rényi and stochastic block model simulations and validates its practical utility on real-world data.