🤖 AI Summary
This study addresses the issue of network autocorrelation induced by node proximity in network data, which frequently leads conventional matching methods to yield spurious causal associations. To mitigate this, we propose a network-constrained matching algorithm that incorporates network proximity as a constraint to optimize the matched set. Furthermore, we develop an asymptotically normal randomization inference procedure that tests the null hypothesis without requiring prespecified correlation structures. An accompanying R package, netmatchRI, is also provided. Simulation studies and an empirical application using the Framingham Heart Study demonstrate that the proposed method effectively eliminates confounding from network autocorrelation and substantially reduces the risk of spurious associations while preserving the validity of causal comparisons, thereby offering a reliable tool for causal inference in networked environments.
📝 Abstract
Matching is widely used to mimic randomized experiments by forming matched sets in which treated and control units differ only randomly with respect to observed covariates. However, when the study population consists of interconnected units from a single network or a small number of networks, matching solely on observed covariates may produce matched units that are more closely connected in the network than would occur by chance. This increased network proximity within matched sets can induce spurious associations between treatment and outcome when both variables exhibit similar autocorrelation patterns on the network. To reduce spurious associations while preserving the validity of causal comparisons, we propose a new matching method that matches units with similar covariates subject to additional network-proximity constraints. For post-matching inference, we propose a randomization-based procedure for testing the sharp null hypothesis of no causal effect. The inference uses the asymptotic normal approximation and accommodates statistical dependence among test statistics obtained from each matched set without requiring explicit specifications of their correlation structures. We demonstrate the validity and utility of the proposed methods through simulation studies and apply them to the Framingham Heart Study. The matching method and subsequent inference procedure are implemented in the R package netmatchRI.