Score
Designs and proves central limit theorems and asymptotic normality results for statistics computed from dependent observations (including network-linked units and cluster-structured data). Builds and analyzes variance formulas and conservative variance estimators, and characterizes how weighting and sampling schemes affect asymptotic behavior to enable valid large-sample inference under cluster or network dependence.
This paper addresses asymptotic inference for a single large network sample generated by strategic interaction and homophily among agents in large-scale static and dynamic network formation models. To overcome the challenge of verifying conventional central limit theorems (CLTs) under network moment dependence, we adapt the exponential stabilization condition from stochastic geometry to network analysis—augmented by branching process theory—to derive verifiable primitive sufficient conditions. The resulting CLT framework requires no repeated sampling and applies directly to a single large network. It substantially broadens the theoretical foundation for network parameter estimation and hypothesis testing, and provides the first asymptotic normality guarantee for strategic network models with explicit, quantifiable regularity conditions.
This paper addresses limit theorems for random variables on network data, overcoming the conventional reliance on Euclidean or metric-space structures. It introduces a weak dependence modeling paradigm applicable to non-embeddable networks—such as financial and social networks—where node positions are unavailable or meaningless. Methodologically, it pioneers the generalization of functional (physical) dependence—originally developed for time series—to arbitrary network topologies, yielding a position-agnostic generalized weak dependence framework. Within this framework, the authors rigorously establish the law of large numbers and the central limit theorem for non-metric networks, and provide verifiable primitive conditions—for instance, dependence decay rates under spatial autoregressive models. The resulting concentration inequalities and limit theory constitute a universal and rigorous foundation for statistical inference on network-structured data.
Gao et al. (JASA 2022) proposed a post-clustering inference framework limited to i.i.d. Gaussian data, failing to accommodate arbitrary dependency structures among observations and features. Method: We generalize their framework to arbitrary dependence settings by developing a unified, dependence-aware post-clustering inference methodology compatible with hierarchical agglomerative clustering (single/complete/average linkage) and k-means. We derive theoretical conditions for well-defined p-values that ensure selective Type I error control and enable consistent covariance matrix estimation. Contribution/Results: Integrating selective inference, high-dimensional statistics, and covariance structure modeling, we design a robust testing pipeline. Experiments on synthetic and real-world protein structural data demonstrate substantial improvements in statistical reliability and practical utility for testing mean differences between clusters under dependence.
Addressing critical limitations of the power-law assumption in modeling degree distributions of complex networks—including substantial bias in parameter estimation, low statistical power in goodness-of-fit (GoF) testing, and neglect of structural heterogeneity—this paper makes three key contributions: (1) a Bayesian framework for unbiased estimation of the power-law exponent α, markedly reducing estimation bias and improving credible interval accuracy; (2) a GoF test based on the Watson statistic, achieving higher statistical power while rigorously controlling Type I error—outperforming the conventional Kolmogorov–Smirnov test; and (3) the first piecewise semiparametric power-law extension model capable of accurately characterizing the *entire* degree distribution, moving beyond traditional tail-only fitting. Extensive simulations and empirical validation on real-world networks (e.g., social and citation networks) confirm that the proposed methods deliver nearly unbiased α estimation, high-power GoF testing, and precise full-distribution modeling—establishing a more robust and comprehensive statistical toolkit for inferring complex network structure.
Existing latent space models for networks primarily focus on point estimation and prediction, lacking a rigorous statistical framework for quantifying estimation uncertainty. Method: We develop the first unified theoretical framework establishing uniform consistency and asymptotic normality of the maximum likelihood estimator under general edge dependence structures and sparse network regimes—applicable to diverse edge types and link functions. Our approach integrates asymptotic statistical theory, uniform convergence analysis, and extensive simulation studies. Contribution/Results: This work extends latent space model inference beyond point estimation to enable principled confidence interval construction and hypothesis testing. It significantly enhances statistical reliability in downstream tasks such as link prediction and network comparison, providing a foundational basis for rigorous statistical inference on network data.
This study addresses limitations of conventional causal inference methods in networked group experiments under interference, which often rely on exposure mapping assumptions, no-interference conditions, or Bernoulli assignment mechanisms. The authors propose a general framework that dispenses with exposure mappings and introduces a novel class of linear weighted estimators tailored to two-stage randomized designs. Certain estimators within this class achieve the optimal root-N convergence rate independent of the number of groups. The work further establishes, for the first time, asymptotic theory and a bias-corrected variance estimation method for dependent statistics under complete randomization. Theoretical analysis confirms the consistency and asymptotic normality of the proposed estimators, while simulations demonstrate their superior finite-sample performance over existing approaches and provide practical guidelines for experimental design and weight selection.
This work addresses the challenges of model selection and hypothesis testing in network data, where strong dependencies and a single observed sample often render existing methods inadequate due to their lack of finite-sample guarantees and limited applicability. The authors propose a general framework based on Universal Inference, which employs edge sampling to partition the network into two subnetworks with controllable dependence structures. This approach yields the first e-value statistic tailored for dependent data, offering strict Type I error control in finite samples and logarithmic consistency under a broad class of alternative models. Empirical evaluations demonstrate that the method achieves both theoretical rigor and strong practical performance in tasks such as random graph model selection and community number estimation.
This study addresses the absence of a statistical inference theory for the Azadkia–Chatterjee conditional dependence coefficient estimator \( T_n \) under general dependence structures. We establish its asymptotic normality and provide, for the first time, a central limit theorem valid for arbitrary dependence. By integrating rank-based and nearest-neighbor constructions, we derive a closed-form expression for the asymptotic variance and propose a consistent variance estimator computable in \( O(n \log n) \) time. Combined with existing bias-correction techniques, our results yield a complete inferential framework for this measure, substantially broadening its applicability in practical data analysis.