Score
Applying statistical techniques (regression adjustment, covariate control, matching) to account for confounding when estimating associations or effects so that observed relationships are tested for robustness to measured covariates.
While covariate-adjusted estimators (e.g., linear regression, matching) are widely used in causal inference, it remains unclear whether rerandomization—despite such adjustments—still delivers meaningful benefits, particularly in finite samples where existing asymptotic theory fails to capture non-precision advantages (e.g., estimator consistency) or practical performance. Method: We conduct large-scale simulation studies to systematically evaluate rerandomization’s impact on estimation precision, statistical power, confidence interval coverage, and consistency across multiple estimators, complemented by theoretical analysis. Contribution/Results: Rerandomization substantially improves finite-sample estimation accuracy, robustness, and consistency of causal effect estimates; enhances statistical power; and reduces false-positive rates—even when covariate adjustment is already employed. These gains extend beyond asymptotic guarantees, offering practitioners a principled, efficiency-enhancing, and reliability-improving design strategy for randomized experiments.
This study addresses a critical yet often overlooked issue in observational research: when proxy variables are used to control for unmeasured confounding, covariates highly correlated with the exposure may inadvertently amplify sensitivity to residual confounding—an effect commonly neglected in conventional sensitivity analyses. Within a regression framework, this work formally characterizes this phenomenon and introduces a novel, observable metric based on the ratio of the exposure model coefficient to the residual variance, which quantifies how covariate structure exacerbates sensitivity to unmeasured confounding. By integrating multicollinearity into the interpretive framework of sensitivity analysis, the approach is validated through linear regression, proxy variable modeling, and sensitivity assessment in the context of smoking and lung cancer. Empirical results demonstrate that increasing socioeconomic stratification over time has heightened the sensitivity of recent data to unmeasured confounding.
Accurate estimation and inference for the average treatment effect (ATE) in multi-covariate randomized controlled trials (RCTs) remain challenging, particularly under high-dimensional covariates and small sample sizes. Method: Building on the Neyman finite-population framework, we propose a bias-corrected regression-adjustment estimator with cross-fitting and introduce, for the first time, an HC3-type heteroskedasticity-robust standard error tailored to high-dimensional settings. We rigorously derive first- and second-order stochastic expansions of the random component of regression-adjustment estimators, identifying the source of higher-order bias in conventional inference; leveraging this insight, we design a cross-fitting procedure to eliminate bias and extend HC3 standard errors to stratified experimental designs. Results: Simulations and reanalysis of Angrist et al. (2009)’s education RCT demonstrate that our method substantially improves estimation accuracy, confidence interval coverage, and statistical power—especially in small-sample and high-dimensional scenarios.
In observational causal inference, weighting methods mitigate covariate imbalance but often inflate variance estimates and yield overly conservative standard errors. This paper proposes augmenting weighted regression with main effects of covariates and their interactions with the treatment variable, integrated with residualization and parametric model augmentation to form a unified inferential framework. We establish, for the first time under design-based, model-based, and finite-sample-corrected superpopulation sampling assumptions, that this approach yields asymptotically valid and more precise standard errors. Theory, simulations, and multiple empirical applications demonstrate substantially narrower confidence intervals—on average 15–30% shorter—with improved inferential accuracy and robustness to both exact and approximately balanced weights. The key innovation lies in achieving simultaneous gains in statistical efficiency and asymptotic validity at minimal variance cost.
This paper addresses bias in causal effect identification arising from unmeasured confounding in observational data. We propose a quantitative sensitivity analysis method grounded in a multiple regression framework. Our key innovation extends the confounding interval approach—previously limited to single-regression settings—to multivariate regression, leveraging observed covariates and domain knowledge (particularly the coefficient of determination, $R^2$) to derive theoretical bounds on omitted-variable bias and thereby achieve partial identification of causal effects. The method supports $R^2$-based bound sensitivity analysis, enabling quantification of estimation uncertainty induced by unmeasured confounders, and is accompanied by an open-source implementation. Simulation studies and empirical applications demonstrate its robustness even under natural stochasticity, offering an interpretable and actionable tool for uncertainty assessment in causal inference.
This paper addresses the challenge of confounder selection in observational studies by proposing an interactive, iterative method that requires neither a pre-specified causal graph nor a complete set of candidate variables. Grounded in latent projection theory, the method dynamically expands a causal graph through successive user-provided local adjustment sets and automatically identifies a minimal “principal adjustment set,” thereby determining whether confounding is controllable. Its key contributions are threefold: (1) it is the first approach to achieve sound and complete confounding control assessment without prior structural assumptions on the causal graph; (2) it makes no assumptions about causal relationships among potential confounders; and (3) it bridges theoretical rigor with practical feasibility. Both theoretical analysis and empirical evaluation demonstrate that, under correct user feedback, the algorithm accurately identifies admissible adjustment sets and correctly determines confounding controllability.
This study addresses a key challenge in randomized controlled trials: how to effectively leverage covariate adjustment to improve the precision of average treatment effect estimation while satisfying regulatory requirements and ensuring statistical validity. The authors propose a prespecified, transparent, and reproducible covariate adjustment framework that, for the first time, integrates data-adaptive methods and machine learning into a regulatory-compliant analytical pipeline. By combining model-misspecification-robust estimation with semiparametric efficiency theory, the approach consistently outperforms unadjusted analyses without compromising causal interpretability or statistical validity. It substantially enhances estimation precision, increases statistical power, and yields narrower confidence intervals.
Measurement error in covariates is pervasive in epidemiology and can induce substantial bias in estimated exposure–outcome relationships, particularly when these associations are nonlinear; yet systematic strategies for correction remain limited. This study presents the first comprehensive evaluation—via blinded, multi-stage simulations—of the performance of six correction methods (pointwise and coefficient-level SIMEX, Bayesian inference, multiple imputation, and regression calibration) combined with four flexible modeling techniques (B-splines, penalized splines, fractional polynomials, and natural splines). Results demonstrate that pointwise SIMEX yields the most accurate and robust estimates overall, while penalized splines, fractional polynomials, and natural splines perform comparably and outperform B-splines. No single approach consistently dominates across all scenarios, underscoring the necessity of conducting sensitivity analyses to account for uncertainty in both measurement error correction and functional form specification.
This study addresses the challenge of reliably estimating the variance of standardized treatment effects in randomized trials with rare binary outcomes or small sample sizes, where existing methods often inflate Type I error rates. The authors propose an influence function–based leave-one-out cross-validation (IF-LOO) variance estimator within the g-computation framework for covariate adjustment. This approach provides, for the first time, a closed-form variance estimator for the standardized average treatment effect that exhibits favorable finite-sample properties, combining computational efficiency with theoretical rigor. Simulation studies demonstrate that IF-LOO effectively controls Type I error in settings with rare events and limited sample sizes, substantially outperforming current methods while remaining readily implementable in clinical trial statistical practice.
This study addresses the limitation of existing E-value methods, which are restricted to single-time-point exposure–outcome relationships and cannot adequately assess the robustness of causal estimates in longitudinal settings with time-varying treatments and confounders. The authors extend the E-value framework to accommodate time-varying confounding by introducing a multi-time-point joint bias factor and propose three sensitivity analysis scenarios: equal-strength distribution, single-time-point dominance, and full-combination visualization, integrated with hazard ratio correction for quantifying causal effect robustness. Simulations reveal that an observed hazard ratio of 1.73 can be nullified by unmeasured confounding associated with the exposure and outcome by as little as 1.96-fold at each time point (single-time-point E-value = 2.85). In a reanalysis of insulin resistance and cardiovascular disease, the time-varying E-value dropped to 1.63 from 2.09, indicating greater sensitivity to unmeasured confounding in longitudinal studies while preserving methodological simplicity and minimal assumptions.
This study addresses the pervasive issue of measurement error in both outcome variables and multiple covariates within routinely collected biomedical data, such as electronic health records, which, if uncorrected, can induce analytical bias and misinform clinical decisions. For the first time within a tutorial framework, it systematically reviews and empirically compares several methods capable of simultaneously correcting measurement error in both outcomes and multiple covariates—including regression calibration, SIMEX, instrumental variable approaches, and modeling strategies leveraging validation subsamples. Through a unified illustrative example and publicly available code, the work not only clarifies the relative performance of these methods in real-world data to guide researchers’ methodological choices but also establishes a reproducible end-to-end analytical pipeline and highlights promising directions for future research.