🤖 AI Summary
This study addresses two critical pitfalls in health digital twins used for clinical decision-making: the fidelity trap—confusing predictive accuracy with causal reasoning—and the feedback trap—where models accumulate bias by updating on data generated from their own recommendations. The work systematically identifies these issues for the first time and proposes a novel design paradigm centered on causal validity, modular architecture, and controlled evolution. By integrating causal inference, module isolation, counterfactual reasoning, and evolutionary governance mechanisms, the framework ensures model reliability in dynamic intervention settings. The research advocates evaluating digital twins based on decision quality rather than mere predictive fit, thereby establishing a theoretical foundation for building trustworthy health digital twins.
📝 Abstract
Digital twins for health may be used to compare treatments, project patient trajectories, and support clinical decisions. While related to mechanical digital twins, those initially developed for engineering applications, replicating the mechanical digital twin architecture and goals may fail in health for two reasons. The fidelity trap is the belief that an accurate model can answer what-if questions by virtue of its accuracy. Prediction and counterfactual reasoning are different tasks, and a twin that can fit past trajectories well may miss the mark when ranking treatments. The feedback trap arises when the twin updates on data its own recommendations helped generate. Refitting in this way can recover a biased relationship and grow more confident even as data grows thinner. We contend that health digital twins should be conceived as causally valid, modular, and evolving systems. Modularity isolates the data and models needed for interventional recommendations, causal validity supports such claims, and governed evolution updates the twin while accounting for how its recommendations reshape the data. We conclude that the standard for a health twin should be how well it supports decisions in the world it helps create, not how faithfully it reproduces the world it observes.