🤖 AI Summary
This study addresses the inflated out-of-distribution (OOD) detection performance in existing methods, which stems from models fitting dataset identity rather than genuine novelty. To resolve this, we introduce a novel whole-dataset hold-out protocol that decouples identity bias from novelty bias. By integrating posterior detection, feature-space directional fitting, and closed-form mathematical derivations, we quantify the inflation of reported predictive gains. Our analysis reveals that conventionally reported improvements are largely spurious and that model capacity is not the primary contributing factor. Furthermore, we establish that a single constant baseline serves as an upper bound for genuine performance gains. This baseline remains robust on a predefined validation set, effectively recalibrating the evaluation standards within the field.
📝 Abstract
A post-hoc out-of-distribution (OOD) detector reads the activations of a trained classifier and returns a score. It fits that score on in-distribution data, and the benchmarks that evaluate it supply a second piece of OOD data for the fitting itself. Some detectors tune a constant on it. Others fit a direction in feature space or train a flexible combiner and report the number that it reaches as the gain that is still available. Every such fit is validated on held-out samples of the same OOD dataset. That check rules out memorizing individual images. It says nothing about a fit that has instead learned which dataset it is looking at, and a direction that recognizes one OOD dataset rather than novelty passes it perfectly. The detector that a practitioner installs meets OOD data from a source that nobody fitted it on, so the difference decides what the reported number is worth. We measure it by holding out the whole OOD dataset rather than a sample of it, and we call that gap the inflation. We read it across a range of combiners on ImageNet and CIFAR-100 backbones. Most of the gain that the usual protocol reports turns out to be dataset identity rather than novelty. The size of the fit does not move what survives, so the effect is not ordinary overfitting. The share depends instead on whether the input exposes class identity, and two controls that vary that property alone separate the inflation on every backbone of both benchmarks. A closed form accounts for the effect and computes it from the fitting rows, so a practitioner can tell which fits will inflate without running the hold-out protocol. One of these fits survives, namely the single constant that the field already picks on a designated validation dataset. Anything above it reports a gain that the hold-out protocol does not return, and on one benchmark what survives falls while what is reported climbs.