🤖 AI Summary
Quantifying classifier robustness under distributional shift and few-shot settings remains challenging due to the lack of reliable, sample-efficient reliability measures. Method: This paper proposes a prediction-level robustness assessment framework for generative probabilistic classifiers, introducing imprecise probability theory—previously unexplored in individual prediction reliability modeling—to jointly characterize generative uncertainty and robust decision-making. Unlike conventional uncertainty quantification paradigms, our approach does not rely on large-sample assumptions and yields verifiable confidence bounds even under scarce training data or test-time distribution shift. Contribution/Results: Empirical evaluation across multiple few-shot and distribution-shift benchmarks demonstrates significant improvements over standard uncertainty baselines. The method provides theoretically grounded, actionable reliability guarantees for high-stakes classification tasks, establishing a novel paradigm for trustworthy classification under limited data and non-stationary environments.
📝 Abstract
Based on existing ideas in the field of imprecise probabilities, we present a new approach for assessing the reliability of the individual predictions of a generative probabilistic classifier. We call this approach robustness quantification, compare it to uncertainty quantification, and demonstrate that it continues to work well even for classifiers that are learned from small training sets that are sampled from a shifted distribution.