Towards Best Practices for Covariate Adjustment in Regulatory Trials: From Fixed to Data-Adaptive Approaches

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses a key challenge in randomized controlled trials: how to effectively leverage covariate adjustment to improve the precision of average treatment effect estimation while satisfying regulatory requirements and ensuring statistical validity. The authors propose a prespecified, transparent, and reproducible covariate adjustment framework that, for the first time, integrates data-adaptive methods and machine learning into a regulatory-compliant analytical pipeline. By combining model-misspecification-robust estimation with semiparametric efficiency theory, the approach consistently outperforms unadjusted analyses without compromising causal interpretability or statistical validity. It substantially enhances estimation precision, increases statistical power, and yields narrower confidence intervals.
📝 Abstract
While randomization justifies the use of unadjusted effect estimators in randomized trials, there is growing interest in covariate adjustment to improve precision. Adjusting for baseline variables that are prognostic of the outcome can reduce estimator variance, resulting in narrower confidence intervals and increased statistical power. Recent guidance by the U.S. Food and Drug Administration supports fixed adjustment for prognostic covariates using parametric regression models. However, this guidance does not address more flexible approaches using data-adaptive or machine learning methods. We offer our perspectives on covariate adjustment to improve analytic precision. We focus on estimating the average effect for the target population in trials with minimal outcome missingness. We provide a non-technical overview of effect estimators that are unadjusted and effect estimators using fixed versus data-adaptive adjustment. We offer practical suggestions for conducting adjusted analyses that are data-adaptive, fully pre-specified, transparently and reproducibly implemented, robust to model misspecification, and guaranteed to improve precision relative to unadjusted analyses --- all while preserving statistical validity and the causal effect of interest. We hope that sharing our perspectives will foster broader discussion and eventual acceptance of principled, pre-specified, data-adaptive covariate adjustment in randomized trials.
Problem

Research questions and friction points this paper is trying to address.

covariate adjustment
randomized trials
data-adaptive methods
statistical precision
regulatory guidance
Innovation

Methods, ideas, or system contributions that make the work stand out.

covariate adjustment
data-adaptive methods
randomized trials
statistical precision
machine learning
🔎 Similar Papers
No similar papers found.
L
Laura B. Balzer
Division of Biostatistics, University of California Berkeley, Berkeley, CA, USA
L
Lei Nie
Division of Biometrics IV, OB/OTS/CDER/FDA, Silver Spring, Maryland, USA
I
Issa J. Dahabreh
Departments of Epidemiology and Biostatistics, Harvard T.H. Chan School of Public Health and Smith Center for Outcomes Research, Beth Israel Deaconess Hospital, Boston, MA, USA
K
Kajsa Kvist
Novo Nordisk A/S, Denmark
T
Tianyue Zhou
Division of Biostatistics, University of California, Berkeley, Berkeley, CA, USA
D
Demissie Alemayehu
Data Science and Analytics, Pfizer Inc., New York, NY, USA
Larry Han
Larry Han
Assistant Professor of Public Health and Health Sciences, Northeastern University
Causal InferenceFederated LearningSurvival AnalysisInfectious Diseases
Z
Zhiwei Zhang
Biostatistics Innovation Group, Gilead Sciences, Foster City, CA, USA
S
Salina P. Waddy
National Center for the Advancement of Translational Science, National Institutes of Health, Bethesda, Maryland
A
Andrew Mertens
Division of Biostatistics, University of California Berkeley, Berkeley, CA, USA
C
Christian B. Pipper
Biostatistics Methods, Novo Nordisk A/S, Copenhagen, Denmark and Section of Epidemiology, Biostatistics and Biodemography, University of Southern Denmark, Odense, Denmark
K
Ken Wiley, Jr.
National Center for the Advancement of Translational Science, National Institutes of Health, Bethesda, Maryland, USA
M
Margot Yann
Forum for Collaborative Research, School of Public Health, University of California Berkeley, USA
G
Gilmer Valdes
OncoBrain Inc and Moffitt Cancer Center, Tampa, Florida, USA
Xu Shi
Xu Shi
University of Michigan
Electronic Health RecordCausal InferenceNegative ControlMachine Translation
Mark van der Laan
Mark van der Laan
Jiann-Ping Hsu/Karl E. Peace Professor of Biostatistics & Statistics, University of California Berkeley
StatisticsBiostatisticsCausal InferenceMachine LearningComputational Biology
M
Maya Petersen
Division of Biostatistics, University of California Berkeley, Berkeley, CA, USA
Kelly Van Lancker
Kelly Van Lancker
Assistant Professor and Postdoctoral researcher in Biostatistics, Ghent University
Biostatisticsclinical trials