š¤ AI Summary
This study addresses the estimation bias and invalid inference arising from inexact matching of continuous covariates in observational studies. It investigates one-to-one matching methods based on a quadratic distance objective, employing asymptotic theory, numerical simulations, and randomization-based inference techniques. The core contribution lies in rigorously establishing, for the first time, that when the covariate dimension satisfies dā¤3, the matching imbalance is negligible, thereby supporting valid Wald and bootstrap inference, while simultaneously revealing the bias-dominated failure mechanism in high-dimensional settings. Furthermore, this work confirms the unbiasedness of low-dimensional matching designs for estimating quantities such as the average treatment effect. These findings provide a clear theoretical foundation delineating the applicability boundaries of matching methods in causal inference.
š Abstract
One-to-one matching without replacement is a classical approach to constructing comparable treated and control samples in the design of observational studies. It pairs each treated unit with a distinct control while minimizing a covariate distance objective. With continuous covariates, the matched pairs generally remain inexact, which contributes to bias in downstream analysis. Its key theoretical properties, such as the resulting imbalance between matched pairs and when it is negligible to support valid inference, remain unclear. In this paper, we analyze one-to-one matching based on $d$-dimensional, continuous covariates with a quadratic covariate-distance objective. First, we find that when $d\leq 3$, under standard conditions on the propensity score ensuring abundant control samples near each treated sample, the imbalance (difference between within-group averages) is root-$n$ negligible uniformly over the family of smooth functions with a common first- and second-order derivative bound. However, such balance is subject to a dimension restriction, as we construct examples in which the imbalance is root-$n$ non-negligible when $d=4$ and dominates root-$n$ rate when $d>4$. Second, we show that when $d\leq 3$, the matched design allows valid Wald-type and bootstrap inference for the average treatment effect on the treated, distributional treatment effects, and quantile treatment effects. Thus, the same outcome-blind matched design supports various downstream inferences without having to tailor the design to the targets. Finally, paired randomization inference based on the matched design is asymptotically valid in the super-population sense for $d\leq 3$ but can fail when $d=4$. We corroborate the theoretical results with numerical experiments.