🤖 AI Summary
Perturbation-based explanation methods suffer from limited reliability due to miscalibrated probability estimates under input perturbations—a fundamental issue stemming from the lack of uncertainty calibration tailored to explainability scenarios. This work establishes, for the first time, a strong empirical and theoretical link between uncertainty calibration and perturbation explanation quality. We propose ReCalX, a post-hoc calibration framework specifically designed for explainability-oriented perturbations. Without altering the original model’s predictions, ReCalX recalibrates confidence distributions over perturbed inputs via an explanation-aware uncertainty alignment mechanism. Integrating principles from probabilistic calibration theory with perturbation sensitivity analysis, ReCalX significantly improves alignment between explanations and human cognition as well as ground-truth object locations. Extensive experiments across multiple benchmarks demonstrate consistent improvements in explanation credibility, stability, and standard interpretability evaluation metrics.
📝 Abstract
Perturbation-based explanations are widely utilized to enhance the transparency of modern machine-learning models. However, their reliability is often compromised by the unknown model behavior under the specific perturbations used. This paper investigates the relationship between uncertainty calibration - the alignment of model confidence with actual accuracy - and perturbation-based explanations. We show that models frequently produce unreliable probability estimates when subjected to explainability-specific perturbations and theoretically prove that this directly undermines explanation quality. To address this, we introduce ReCalX, a novel approach to recalibrate models for improved perturbation-based explanations while preserving their original predictions. Experiments on popular computer vision models demonstrate that our calibration strategy produces explanations that are more aligned with human perception and actual object locations.