🤖 AI Summary
This study addresses the overconfidence and uncertainty estimation failures of evidential deep learning under adversarial and out-of-distribution (OOD) inputs. To this end, it proposes CLEAR, a lightweight post-hoc calibration framework that requires no retraining. This method introduces a novel task-agnostic latent consistency mechanism that leverages calibration data to characterize the geometric structure of the latent space, dynamically rectifying evidence strength through perturbed view generation and conflict measurement. Consequently, CLEAR significantly enhances anomalous input detection without altering base predictions. Evaluated on the ImageNet-to-CUB shift, it improves OOD and adversarial AUROC by 8.29% and 5.01%, respectively, while achieving 17.4× faster inference than comparable methods and preserving multi-task performance.
📝 Abstract
Reliable uncertainty quantification is essential for deploying deep learning models in high-stakes settings, where out-of-distribution and adversarial inputs can induce confident but unreliable predictions. Evidential Deep Learning provides efficient uncertainty estimates in a single forward pass, but can still assign high evidential strength to inputs that are poorly supported by the learned representation, such as adversarial inputs. We introduce CLEAR, a lightweight, task-agnostic post-hoc method that improves evidential robustness without retraining or altering the base prediction. Using held-out calibration data, CLEAR characterises the group-conditioned geometry of the model's latent space. At inference, it efficiently generates perturbation views directly in the latent space and measures their conflict relative to the calibrated geometry of the predicted group. High latent conflict indicates unsupported evidence, which CLEAR uses to selectively reduce evidential strength while retaining evidence for latent-consistent inputs. On ImageNet$\rightarrow$CUB, CLEAR improves OOD and adversarial AUROC by $+8.29$ and $+5.01$ while running 17.4$\times$ faster than competing post-hoc methods while preserving predictive performance across classification, regression, and object detection benchmarks.