๐ค AI Summary
Although learning models with noisy labels perform well on closed-set tasks, they suffer from โuncertainty collapse,โ wherein misclassified in-distribution samples and out-of-distribution (OOD) samples become indistinguishable due to overlapping features and confidence scores. This work is the first to identify this phenomenon and introduces ACC-OOD, a learner-agnostic benchmark that uniformly evaluates a modelโs ability to detect both near- and far-OOD samples. To mitigate uncertainty collapse, the authors propose Virtual Margin Regularization (VMR), a lightweight method that enhances OOD separability without compromising closed-set accuracy. Extensive experiments demonstrate that VMR significantly improves OOD detection performance while maintaining high in-distribution classification accuracy, underscoring the necessity of jointly evaluating noise robustness and open-world reliability.
๐ Abstract
Learning with noisy labels (LNL) is typically benchmarked by closed-set classification accuracy, yet deployment often requires classifiers to reject out-of-distribution (OOD) inputs. We present a learner-agnostic ACC-OOD benchmark that freezes LNL checkpoints and evaluates them with standardized near-/far-OOD routing and post-hoc scores across synthetic and real label noise. The benchmark reveals a recurring failure mode: high closed-set accuracy does not ensure OOD reliability, because low-confidence, misclassified in-distribution samples can overlap the score and feature regions occupied by OOD inputs under noisy training. We term this pathology uncertainty collapse. This structural overlap can make high-accuracy LNL methods lose separability at the ID-error/OOD interface under standard OOD scores. As an intervention, we study Virtual Margin Regularization (VMR), a lightweight repair probe demonstrated mainly with PSSCL that synthesizes boundary virtual outliers on trusted ID batches and widens the energy margin. VMR partially reduces the collapse-induced far-OOD failure without replacing the host objective or sacrificing closed-set accuracy in the tested settings. These results support LNL benchmarks that co-report closed-set generalization, open-world reliability, and structural overlap diagnostics.