🤖 AI Summary
This work addresses the limitation of existing methods that can either detect out-of-distribution samples or quantify uncertainty but struggle to pinpoint the specific causes of model failure. The authors propose a self-diagnosing model that jointly learns structured failure attribution signals alongside its primary predictions, extending scalar uncertainty into an attribution vector capable of distinguishing among four distinct failure modes: covariate shift, semantic shift, noisy corruption, and adversarial perturbations. The attribution vector is generated by a neural network and constrained via a consistency regularization term that aligns uncertainty estimates with attribution predictions. To evaluate the approach, the authors construct a benchmark dataset incorporating predefined shift mechanisms. Experimental results demonstrate that the method not only effectively detects anomalies but also accurately attributes their underlying failure types, significantly enhancing model interpretability and robustness.
📝 Abstract
Distribution shift poses a significant challenge to the robustness of machine learning models, but the current solutions only aim to detect out-of-distribution (OOD) samples and predict uncertainty levels. We introduce a problem setting for failure attribution under distribution shift, which enables the models not only to detect OOD samples, but also to find out the reason for their failure. The solution we propose is called self-diagnosing models, which are capable of jointly learning predictive output, predictive uncertainty, and a failure attribution signal. In particular, we use the failure attribution vector, produced by a neural network, which provides a structured representation of predictive unreliability by distinguishing four different types of failures: covariance shift, semantic shift, noise corruption, and adversarial perturbation. In other words, we move from scalar uncertainty towards failure identification. For training the model, we introduce a consistency regularizer that encourages consistency between uncertainty and failure attribution predictions. Moreover, to be able to evaluate the model on its ability to find the reasons for failure, we construct several distribution shift benchmarks with predefined mechanisms for generating distribution shifts.