g-computation

Design and implement g-computation procedures that use a specified causal model to simulate counterfactual (potential) outcomes under hypothetical interventions and aggregate those simulated outcomes to produce causal effect estimates (e.g., average potential outcomes, contrasts such as risk differences or ratios). Produce valid uncertainty quantification for those estimates by deriving or approximating standard errors and confidence intervals (via analytical influence-function expressions, asymptotic theory, bootstrap, or Monte Carlo sampling) and implement the simulation and aggregation steps required to obtain point estimates and interval estimates.

g-computation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.17
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Constructing g-computation estimators: two case studies in selection bias

Jun 03, 2025
PZ
P. Zivich
🏛️ University of North Carolina at Chapel Hill | Yale University

This paper addresses selection bias in epidemiology under complex causal structures—specifically, the coexistence of treatment-induced selection and multiple biases (e.g., absence of a joint adjustment set). We propose two novel g-computation estimators. Methodologically, we are the first to unify these two challenging forms of selection bias within the g-computation framework, implementing them via stacked estimating equations. Leveraging causal graph identification, Monte Carlo simulation, and finite-sample theoretical analysis, we rigorously establish consistency and asymptotic normality. Simulations demonstrate that the new estimators significantly outperform conventional adjustment methods in small samples. Our core contribution lies in bridging causal identification with implementable estimation: systematically translating causal graph–based inferential results into interpretable, computationally feasible, and statistically guaranteed estimation strategies.

Adapting g-computation to address complex selection biasesImplementing g-computation for treatment-induced selection biasTranslating causal diagrams into novel estimation strategies

Bias reduction in g-computation for covariate adjustment in randomized clinical trials

Sep 08, 2025
XZ
Xin Zhang
🏛️ Pfizer Inc. | Institute of Natural Sciences | CMA-Shanghai | SJTU-Yale Joint Center for Biostatistics and Data Science | Shanghai Jiao Tong University | Division of Biostatistics and Health Data Science | University of Minnesota Twin Cities

In randomized clinical trials, g-computation often yields biased treatment effect estimates and underestimates variance after covariate adjustment—particularly in small samples or rare-event settings—where maximum likelihood estimation may fail. To address this, we introduce, for the first time, a systematic bias-correction framework into g-computation. Our method employs a generalized Oaxaca–Blinder estimator for debiasing, integrated with Firth’s penalized likelihood correction and asymptotic bias analysis to derive a bounded, robust variance adjustment. The resulting estimator improves finite-sample accuracy and inferential stability without compromising efficiency. Through extensive simulations and reanalyses of real clinical trials, we demonstrate that our approach effectively balances the bias–efficiency trade-off, yielding more reliable and practically applicable unconditional treatment effect estimates.

Addressing small sample and rare event estimation issuesImproving finite-sample performance and inference validityReducing bias in g-computation for randomized trials

Integer Programming for Generalized Causal Bootstrap Designs

Oct 28, 2024
JB
Jennifer Brennan
🏛️ Google Research | University of Southern California | Uber Technologies

In experimental causal inference, design uncertainty—arising from the assignment mechanism—is often overlooked under small sample sizes and heterogeneous treatment effects; conventional causal bootstrap methods apply only to completely randomized designs and average treatment effect estimation. This paper addresses this limitation by introducing integer linear programming into the causal bootstrap framework for the first time, enabling computation of the worst-case copula under generalized assignment mechanisms (e.g., conditional unconfoundedness, bounded confounding) to uniformly calibrate design uncertainty. The method accommodates both linear and quadratic treatment effect estimators and is supported by asymptotic theory establishing its validity. Monte Carlo simulations demonstrate that, in small-scale geographic experiments, the proposed approach substantially improves confidence interval coverage and precision while delivering more robust control of Type I error.

Addresses design uncertainty in small fixed heterogeneous samplesExtends causal bootstrap to non-standard designs and estimatorsGeneralizes method for various assignment types and estimators

This study addresses the bias introduced by external control data in hybrid controlled trials and the reliance of conventional methods on strong exchangeability assumptions. The authors propose a model-robust G-computation approach that achieves unbiased and efficient estimation under weaker assumptions by adjusting for baseline covariates. Notably, the method does not require exchangeability and retains consistency and asymptotic normality even when the outcome regression model is misspecified, offering robustness, simplicity, and efficiency. Integrating variable selection with covariate adjustment, theoretical analysis, simulation studies, and an empirical application to an HIV treatment trial demonstrate that the proposed method effectively controls bias and substantially improves estimation efficiency across diverse scenarios.

bias mitigationexchangeabilityg-computation

This study addresses the bias in causal effect estimation arising from the coexistence of unmeasured cluster-level confounding and treatment effect heterogeneity in observational clustered data. To tackle this dual challenge, the authors propose an intra-group g-computation approach: clusters are first stratified by observed treatment prevalence, within-stratum g-computation is implemented using random-effects models, and estimates across strata are then aggregated to correct for bias. This method innovatively embeds random-effects modeling within the g-computation framework, effectively mitigating both sources of bias. Simulation studies demonstrate that the proposed estimator achieves the lowest root mean squared error when unmeasured confounding and heterogeneity co-occur. Applied to data from Bangladesh, the method reveals that adolescent pregnancy is associated with an average reduction of 0.12 in child height-for-age Z-scores (95% CI: [–0.18, –0.06]).

causal effect estimationhierarchical dataobservational studies

Latest Papers

What's happening recently
View more

This study addresses a critical limitation of the traditional synthetic control method: when treatment effects are weak, systematic bias can shift the center of confidence intervals, leading to misleading inferences. To remedy this, the authors propose a novel time placebo–guided approach that explicitly quantifies and corrects for such bias. By retrospectively assigning placebo intervention dates within the observed panel and refitting the synthetic control model at each, the method directly estimates the bias distribution under the null hypothesis. This enables the construction of nonparametric confidence intervals calibrated to maintain nominal coverage regardless of the true effect trajectory. The proposed procedure achieves stable, bias-corrected inference with fixed interval width, substantially enhancing the robustness of causal conclusions in synthetic control applications.

biascausal inferenceconfidence intervals

This study addresses the challenge of accurately estimating the causal effects of multiple interventions on average length of stay in hospital quality improvement, where data scarcity and complex underlying mechanisms hinder reliable inference. The authors propose expert-guided g-computation (egg-computation), a novel framework that integrates Gantt charts with causal directed acyclic graphs (DAGs) to unify expert knowledge and empirical evidence. Clinical expert judgment is selectively incorporated only where causal identification is otherwise impossible, and large language models (LLMs) are leveraged to scalably generate causal graphs and estimates of time savings. In simulations, the method outperforms conventional causal inference approaches; when applied to evaluate eleven real-world hospital interventions, LLM-assisted results show high concordance with human expert assessments, demonstrating an efficient and scalable solution for causal effect estimation.

causal effectsg-computationhospital quality improvement

This study identifies and quantifies a critical implementation flaw in the gsynth R package (prior to version 1.3.1) when combining interactive fixed effects–expectation maximization (IFE-EM) estimators with parametric bootstrap inference: the algorithm erroneously substitutes in-sample residuals for out-of-sample prediction errors, leading to systematically underestimated standard errors and compromised inferential validity. Through Monte Carlo simulations, placebo tests, recomputed standard errors, and robustness checks using the generalized synthetic control method (GSCM), we demonstrate that the original approach yields substantially inflated false positive rates. After correction, most estimated treatment effects lose statistical significance, and reanalysis of three APSR articles using GSCM invalidates their core conclusions, underscoring the essential role of methodological rigor in empirical political science research.

gsynthimplementation errorInteractive Fixed Effects

This study addresses the limited accessibility of the g-formula in causal inference due to its mathematically opaque formulation for those with modest statistical backgrounds. Under the standard assumptions of consistency, positivity, and conditional exchangeability, the authors systematically reformulate the g-formula into two nonparametrically equivalent representations: a non-iterative (NICE) and an iterative (ICE) form. This novel decomposition clarifies the g-formula’s intrinsic connections to the law of iterated expectations and conditional expectation operators. Through three progressively complex numerical examples—spanning settings with both fixed and time-varying confounders—the paper intuitively illustrates the computational mechanics and causal identification logic underlying the g-formula. The proposed framework substantially enhances interpretability and generalizability, offering practitioners a transparent and unified pathway for estimating causal effects.

causal effect identificationcausal inferenceg-formula

This study addresses the challenges of estimating causal effects of time-varying interventions on rare survival outcomes in large-scale longitudinal observational studies, where high computational costs and severe class imbalance often hinder reliable inference. The authors propose a subsampling and inverse probability reweighting framework tailored for longitudinal survival data, which integrates seamlessly with existing causal estimators—such as g-formula–based ICE—while preserving estimator consistency and substantially reducing computational burden. This approach represents the first application of a subsampling strategy that jointly optimizes computational efficiency and statistical consistency in the context of causal inference for longitudinal rare events, effectively mitigating model instability induced by outcome imbalance. Simulations and an empirical analysis using electronic health records to assess the impact of social-behavioral factors on suicide risk demonstrate that the method markedly improves computational efficiency while enhancing both the stability and accuracy of causal estimates.

causal effect estimationclass imbalancecomputational scalability

Hot Scholars

JJ

Julie Josse

Senior Researcher Inria,
Missing valuesLow rank matrixcausal inferenceR
JB

John B. Carlin

Murdoch Childrens Research Institute, University of Melbourne
biostatisticsepidemiology
MM

Margarita Moreno-Betancur

Professor of Biostatistics, University of Melbourne & Murdoch Children's Research
Causal inferenceMissing dataSurvival analysis
NP

Niels Peek

The Healthcare Improvement Studies Institute, University of Cambridge
data sciencehealthcare improvementhealth informaticsartificial intelligence
SV

Stijn Vansteelandt

Professor of Statistics, Ghent University
Causal inferenceCausal Machine LearningEpidemiologic methodsMediation analysis