🤖 AI Summary
This work proposes a general-purpose loss function based on a Bayesian latent-variable switching mixture model that simultaneously enables robust learning and unsupervised identification of corrupted labels during training. While existing robust loss functions mitigate the adverse effects of label noise, they lack the ability to explicitly detect contaminated samples. The proposed approach uniquely unifies robust loss with interpretable posterior probabilities of label corruption, incorporates input-dependent priors to capture the spatial locality of noise, and naturally induces an Occam’s razor regularization via marginal likelihood to prevent over-flagging. Experiments on CIFAR-10 under asymmetric label noise (corruption rates 0.2–0.6) demonstrate that the method accurately recovers the underlying corruption structure, effectively distinguishes clean from noisy samples, identifies the direction of label flips, and significantly outperforms four state-of-the-art robust loss baselines.
📝 Abstract
Engineered robust losses such as Huber, Student-$t$, and generalised cross-entropy make supervised models tolerant of contamination but cannot answer which observations are corrupted. We introduce Neural Bayesian Anomaly Mitigation (NBAM), a general-purpose drop-in loss derived from a Bayesian latent-switch mixture model: the marginal likelihood defines a robust supervised loss, and the associated posterior defines an unsupervised contamination classifier. Like Huber or Student-$t$, NBAM can replace the standard training loss in any supervised pipeline; unlike them, it additionally learns a structured contamination model and returns a calibrated per-sample contamination posterior. A learned input-dependent prior $π_φ(x)$ captures the spatial locality of contamination, so that samples near known corruptions are more likely to be flagged, while an Occam penalty emerges automatically and regularises against over-flagging. On CIFAR-10 with asymmetric label contamination, NBAM recovers the structure of the corruption process without supervision: the contamination posterior separates clean from corrupted samples, and the learned anomaly head identifies the direction of every label-flip pair. Alongside these capabilities, NBAM outperforms the four robust-loss baselines considered here at contamination rates 0.2-0.6.