🤖 AI Summary
This study addresses the vulnerability of confidence calibration in medical image classification models under adversarial attacks, which often leads to calibration failure and a lack of safety guarantees. To mitigate this, we propose a posterior certification strategy that integrates statistical concentration inequalities with local Lipschitz constant estimation to derive worst-case upper bounds on the binned calibration error of arbitrary classifiers within an $\ell_2$-norm ball constraint. This work provides the first adversarial robustness certification specifically targeting confidence calibration rather than merely predicted labels. Extensive evaluations across eleven medical imaging tasks demonstrate that our approach significantly outperforms existing baselines, achieving higher certified calibration coverage while maintaining tight error bounds.
📝 Abstract
Deep neural networks remain vulnerable to adversarial perturbations, which can distort not only predictions but also confidence scores, undermining uncertainty calibration. While existing certification methods focus on preserving the predicted category, providing guarantees on how calibration behaves under adversarial attacks remains overlooked. In this work, we introduce CalCErt, a simple and efficient post-hoc strategy that certifies bin-wise confidence calibration for any pretrained differentiable classifier. Our approach combines empirical calibration estimates, statistical concentration bounds, and local Lipschitz estimates of the confidence function to derive data-dependent upper bounds on worst-case miscalibration within an $ell_2$-ball of radius R. We evaluate CalCErt across 11 medical image classification tasks and multiple adversarial perturbations, demonstrating substantially higher certified coverage than baseline strategies while maintaining competitive tightness. Our code is available at https://github.com/leofillioux/calcert.