🤖 AI Summary
This study addresses the challenge of quantifying the benefit and reliably assessing uncertainty when integrating real-world data with randomized controlled trials. Building on adaptive targeted maximum likelihood estimation (A-TMLE), the authors propose three reproducible tools: a bias model audit report card, a fusion gain efficiency plot, and a selection-aware inference method for data-adaptive gain estimators. Innovatively introducing an auditable bias assessment framework, the work demonstrates that fusion efficiency is governed primarily by the magnitude of bias rather than functional complexity, and establishes a calibrated inference framework. Empirical validation across HIV treatment, public health, and vocational training applications reveals that only the modular jackknife yields calibrated—albeit conservative—confidence intervals, substantially enhancing the reliability of real-world evidence.
📝 Abstract
Augmenting a randomized controlled trial with real-world data promises greater efficiency, but how much a given fusion actually delivers, and how to attach honest uncertainty to that gain, is rarely characterized. Using adaptive targeted maximum likelihood estimation (A-TMLE) as the running example, we develop three reproducible tools for honest evidence from combined trial and real-world data. First, a report card that makes the estimator's data-adaptively learned bias model auditable, measuring how well it recovers the true enrollment-effect surface and attributing the estimator's variance to its structural parts. Second, a map of when fusion helps versus hurts, benchmarked against an efficient trial-only estimator: the gain is driven primarily by the magnitude of the real-world bias rather than its functional complexity, a dominance an exact variance identity explains; it crosses break-even near a moderate bias and erodes as the trial grows, so the advantage is finite-sample rather than a form of super-efficiency. Third, selection-aware inference for the gain, treated as a data-adaptive estimand: the naive standard error undercovers, and among ten candidate standard errors only a block jackknife is calibrated, though conservatively so. Three openly available fusions, in a biomedical HIV trial, a public-health trial, and a job-training trial, span the map and show the difference an honest interval makes for real-world evidence.