🤖 AI Summary
This study addresses the unreliability of out-of-distribution (OOD) detection in semantic segmentation, which frequently stems from model overconfidence and prohibitive computational overhead. To overcome these limitations, this work proposes a Wasserstein distance-based evidential deep learning framework that models Dirichlet distributions via a single forward pass. By replacing conventional Euclidean losses with a Wasserstein objective for optimization, the approach achieves efficiently calibrated uncertainty estimation while respecting the geometric structure of the probability simplex. Furthermore, this research reveals that the optimal Wasserstein order varies systematically across backbone architectures, specifically between CNNs and Transformers. Extensive evaluations on benchmarks such as SegmentMeIfYouCan demonstrate that the proposed method surpasses existing approaches, delivering robust OOD detection with minimal computational cost.
📝 Abstract
Semantic segmentation networks operate on a fixed set of classes and therefore fail when out-of-distribution (OOD) objects appear during deployment, a critical limitation for safety-critical applications such as autonomous driving. Reliably identifying OOD objects requires well-calibrated epistemic uncertainty, yet common softmax-based confidence scores remain overconfident, while Bayesian alternatives such as Monte Carlo dropout or deep ensembles require costly repeated forward passes. Evidential Deep Learning (EDL) offers an efficient alternative by modeling class probabilities as a Dirichlet distribution learned from a single deterministic forward pass. Existing EDL formulations rely on Euclidean objectives that push predictions towards the simplex vertices, encouraging overconfidence rather than preserving uncertainty for unfamiliar inputs. We instead employ Wasserstein-based objectives, which respect the geometry of the probability simplex, and study the influence of the Wasserstein order on segmentation accuracy and OOD detection within a unified evidential framework. We evaluate this framework on a convolutional (DeepLabV3+) and a transformer-based (SegFormer) architecture on the SegmentMeIfYouCan benchmark, including LostAndFound, RoadObstacle21, RoadAnomaly21, and Fishyscapes. Our results show the optimal Wasserstein order is architecture-dependent: second-order objectives dominate on the convolutional backbone, third-order objectives on the transformer backbone, and our framework surpasses comparable baselines on most metrics, with a single deterministic forward pass.