Score
Assessing covariate balance and the credibility of causal comparisons by preparing and validating covariate records, testing sensitivity to confounding and alternative estimators, and choosing experimental or matching designs to rule out selection bias.
While covariate-adjusted estimators (e.g., linear regression, matching) are widely used in causal inference, it remains unclear whether rerandomization—despite such adjustments—still delivers meaningful benefits, particularly in finite samples where existing asymptotic theory fails to capture non-precision advantages (e.g., estimator consistency) or practical performance. Method: We conduct large-scale simulation studies to systematically evaluate rerandomization’s impact on estimation precision, statistical power, confidence interval coverage, and consistency across multiple estimators, complemented by theoretical analysis. Contribution/Results: Rerandomization substantially improves finite-sample estimation accuracy, robustness, and consistency of causal effect estimates; enhances statistical power; and reduces false-positive rates—even when covariate adjustment is already employed. These gains extend beyond asymptotic guarantees, offering practitioners a principled, efficiency-enhancing, and reliability-improving design strategy for randomized experiments.
In causal inference, conventional covariate balance tests suffer from inflated false-positive rates when irrelevant covariates are imbalanced and exhibit low sensitivity to imbalance in potential outcomes. To address these limitations, we propose a conditional balance test grounded in prognostic covariate importance—explicitly incorporating each covariate’s predictive strength for potential outcomes into the test weighting scheme. This enables joint optimization of statistical power and false-positive control. Our method employs a standardized regression-weighted mean difference test, supported by theory-driven weight construction and a Monte Carlo simulation validation framework. We provide theoretical guarantees of improved statistical power. Simulation studies demonstrate that our approach achieves substantially higher detection power than global balance tests under potential outcome imbalance, while reducing the false rejection rate due to irrelevant covariate imbalance by over 40%.
This study addresses covariate imbalance arising from limited sample sizes in survey experiments by proposing an integrated experimental design that combines stratified rejection sampling with rerandomization, coupled with a covariate-adjusted analysis approach. The framework enhances covariate balance during the design phase and improves estimation efficiency during the analysis phase. Theoretically, the authors derive the design-based asymptotic distribution of the stratified difference-in-means estimator, demonstrating that its limiting distribution is more concentrated around the true average treatment effect than those of existing methods. Numerical experiments confirm that the proposed method substantially increases precision and efficiency while preserving unbiasedness of the treatment effect estimates.
This paper investigates how the “design phase”—i.e., subsample selection to achieve covariate balance between treatment and control groups—affects causal inference via linear regression in observational studies. Methodologically, it formalizes subsample selection as an estimator adjustment process centered on covariate balancing, rigorously establishing its theoretical role in mitigating bias from model misspecification. It further introduces a sensitivity analysis framework grounded in imbalance metrics, serving both as a quantitative measure of design quality and a transparency vehicle for results. The key contribution lies in unifying the design and estimation phases within the linear regression framework for the first time, thereby elevating covariate balance from a heuristic practice to a theoretically grounded principle and operational standard for bias control. This integration substantially enhances the robustness and reproducibility of causal inference.
This paper addresses the estimation of the average treatment effect on the treated (ATT) in difference-in-differences (DID) designs with panel data. We propose a novel ATT estimator that integrates covariate-balancing propensity scores (CBPS) into the DID framework, achieving global covariate balance between treatment and control groups via optimized weighting. The estimator combines local efficiency with double robustness. We establish its theoretical superiority: under local misspecification of either the propensity score or outcome model, it converges faster than augmented inverse probability weighting (AIPW). Simulation studies and empirical applications demonstrate that the estimator exhibits smaller finite-sample bias, higher confidence interval coverage, and more stable inference—substantially enhancing the reliability and practical applicability of DID estimation.
This study addresses the lack of large-scale empirical comparisons among covariate adjustment strategies in randomized clinical trials (RCTs), which has led to ambiguity in method selection and covariate inclusion criteria. Leveraging individual participant data from 50 publicly available RCTs (29,094 participants; 574 treatment–outcome pairs), we systematically evaluate the statistical efficiency gains of 18 adjustment strategies—spanning classical regression, inverse probability weighting, and machine learning estimators—combined with three covariate selection rules. Results show that covariate adjustment reduces variance by an average of 13.3% for continuous outcomes and 4.6% for binary outcomes. Notably, parsimonious and transparent methods such as analysis of covariance consistently outperform default configurations of complex machine learning models, providing the first large-scale empirical evidence supporting their preferential use in routine RCT analyses.
In causal inference, the consistency assumption—that potential outcomes under treatment coincide with observed outcomes, and that there are no hidden versions of treatment—is often violated, particularly due to unobserved treatment heterogeneity (e.g., variation in surgeons’ skill levels), rather than classical unmeasured confounding. Method: We formally distinguish treatment versions from covariates and develop the first sensitivity analysis framework specifically tailored to hidden treatment versions. Leveraging a novel notation system, we integrate causal diagrams with counterfactual models to define parameterized sensitivity measures and derive bounds on causal effects. Contribution/Results: Our framework is empirically validated in real-world surgical settings. It quantifies how violations of consistency bias causal effect estimates and substantially enhances the rigor of robustness assessment for causal conclusions—enabling principled evaluation of estimation uncertainty arising from treatment-version heterogeneity.
Estimating the functional relationship between a continuous exposure and a binary outcome is challenging when covariates are measured with error. This study presents the first systematic evaluation of Simulation-Extrapolation, Regression Calibration, multiple imputation, and Bayesian correction methods, each coupled with flexible modeling techniques—including B-splines, P-splines, and fractional polynomials—within a multi-team, fully blinded, neutral simulation framework. By generating 155 distinct simulation scenarios and repeated samples, the research quantifies the bias and variance of each approach, revealing their relative strengths and limitations. The findings not only inform method selection under measurement error but also demonstrate the feasibility and value of this neutral comparative paradigm for rigorous methodological assessment.
Existing data-adaptive methods for identifying treatment effect modifiers often lack finite-sample control of false positives, leading to irreproducible findings. This work proposes a general framework that integrates cross-fitted conditional average treatment effect estimation with ensemble pathwise stability selection, accommodating arbitrary treatment effect estimators and variable selectors. For the first time, it provides a non-asymptotic upper bound on the expected number of false positives in modifier selection under finite samples. The method establishes a theoretical link between the accuracy of treatment effect estimation and the performance of variable discovery. Empirical validation on both a randomized oncology trial and an observational study of maternal smoking’s effect on infant birth weight demonstrates its ability to control the false discovery rate while achieving strong selection consistency.
This study addresses the lack of systematic criteria for covariate selection in difference-in-differences (DID) designs, which often leads to arbitrary specifications of the conditional parallel trends assumption. It introduces causal graphical models to formally identify the set of covariates required to satisfy this assumption, highlighting the critical role of time-invariant covariates and distinguishing between treatment types and implementation strategies to clarify treatment–confounding feedback mechanisms. By integrating a multi-period DID framework with conditional independence tests and consistency analysis of adjustment sets, the paper proposes principled criteria for covariate inclusion, resolving the prevalent mismatch between adjustment sets and estimation methods. This approach substantially enhances the identification validity and robustness of causal effect estimates in DID analyses.