🤖 AI Summary
This study addresses the performance degradation of clinical AI models across institutions due to data distribution shifts, where retraining is constrained by scarce annotations and regulatory requirements. To overcome this, we propose GUARD, a framework that updates deployed models to resist distribution drift via semi-supervised adversarial optimization, leveraging limited target labels alongside multi-source heterogeneous data. Methodologically, GUARD innovatively integrates double machine learning debiased estimation, cross-fitting, and uncertainty-aware Group Distributionally Robust Optimization (Group DRO) to enable efficient semi-supervised transfer while ensuring valid statistical inference. Evaluated on rheumatoid arthritis prediction tasks using electronic health records, the framework sustains high predictive accuracy over multiple years relying solely on minimal yearly annotations, significantly outperforming conventional transfer learning approaches.
📝 Abstract
Artificial intelligence models deployed as clinical decision support tools often suffer substantial performance degradation over time and across institutions due to covariate shift, concept drift, and cross-system heterogeneity. Fully retraining complex models is frequently infeasible, particularly in EHR settings where labeled outcome data are scarce and regulatory constraints limit model modification. We propose GUARD (Guided and Uncertainty-Aware Robustness to Domain shift), a unified statistical framework for updating an existing deployed model using limited labeled target data and multiple heterogeneous source populations while guarding against future distributional shifts. GUARD formulates robust multi-source transfer learning as a semi-supervised, adversarial optimization problem anchored at the current model, and employs double machine learning with cross-fitting to obtain debiased, efficient estimates that leverage abundant unlabeled target covariates. It then constructs an uncertainty-aware, guided group DRO estimator that combines source and target information while accounting for sampling variability in the source models. Our framework jointly provides principled robustness to domain shift, efficient semi-supervised estimation, and valid statistical inference for multi-source transfer of clinical prediction models. We demonstrate the utility of GUARD via extensive simulation experiments and a real world application to predicting future disease activity in rheumatoid arthritis from EHR data spanning more than a decade. GUARD recalibrates a pretrained model with only a small number of labeled records per year and remains accurate over multi-year horizons where target-only and standard transfer estimators degrade sharply.