🤖 AI Summary
This study addresses the challenge of poorly calibrated prediction probabilities in black-box radiology AI models, where internal parameters are inaccessible, thereby compromising clinical decision safety. To overcome this limitation, this work proposes DualTTA, a model-agnostic test-time augmentation framework that requires neither internal model information nor training data. DualTTA pioneers the generation of test samples through 3D CT geometric and physical perturbations and introduces a probability-level learned aggregation strategy to optimize predictive confidence. Evaluated on pulmonary embolism and intracranial hemorrhage detection tasks, the proposed method significantly reduces expected calibration error by 54% and 43%, respectively. Notably, it outperforms several conventional calibration approaches that rely on internal model access, offering an efficient solution for the safe deployment of black-box medical AI systems.
📝 Abstract
Radiology AI systems increasingly inform clinical decisions such as triage, follow-up imaging, and treatment planning. For these decisions to be made safely, model outputs must be well calibrated, meaning predicted probabilities accurately reflect true risk. Many standard techniques for improving calibration, such as MC Dropout and Deep Ensembles, require access to model parameters or retraining. However, proprietary clinical AI systems operate as black boxes, preventing access to the model's internals. To that end, we propose a model-agnostic framework for improving calibration of black-box models using clinically grounded test-time augmentation (TTA). Our framework applies geometric and physics-inspired 3D CT perturbations and learns probability-level aggregation strategies without access to model internals or the original training data. Across pulmonary embolism and intracranial hemorrhage detection tasks, DualTTA achieved the strongest overall calibration among TTA methods, reducing the Expected Calibration Error by 54% (0.239 -> 0.109) and 43% (0.051 -> 0.029), respectively, while requiring only input-output access. Additionally, DualTTA outperformed uncertainty estimation techniques that require access to model internals, such as Temperature Scaling, MC Dropout, and Deep Ensembles, in most calibration metrics. These results demonstrate that learned TTA aggregation can improve the calibration of clinical AI systems, providing a practical approach for improving the reliability of black-box medical AI.