🤖 AI Summary
This work addresses the challenge of effectively evaluating the reliability of individual predictions made by classifiers under distribution shift. We systematically compare robustness quantification (RQ) and uncertainty quantification (UQ) for this task and, for the first time, explore the potential of integrating both approaches. Through extensive experiments across multiple benchmark datasets—covering both standard training conditions and various distribution shift scenarios—we elucidate the conceptual distinctions between RQ and UQ. Our results demonstrate that RQ consistently matches or outperforms UQ in most settings, and that hybrid strategies combining RQ and UQ significantly enhance the accuracy of reliability assessment for individual predictions. These findings offer a novel perspective toward building more trustworthy AI systems.
📝 Abstract
We consider two approaches for assessing the reliability of the individual predictions of a classifier: Robustness Quantification (RQ) and Uncertainty Quantification (UQ). We explain the conceptual differences between the two approaches, compare both approaches on a number of benchmark datasets and show that RQ is capable of outperforming UQ, both in a standard setting and in the presence of distribution shift. Beside showing that RQ can be competitive with UQ, we also demonstrate the complementarity of RQ and UQ by showing that a combination of both approaches can lead to even better reliability assessments.