Class Imbalance in Anomaly Detection: Learning from an Exactly Solvable Model

📅 2025-01-20
📈 Citations: 0
Influential: 0
📄 PDF

career value

231K/year
🤖 AI Summary
Class imbalance severely degrades anomaly detection performance, yet its theoretical underpinnings remain poorly understood. Method: We establish a rigorous teacher–student perceptron framework and, for the first time, apply replica theory to derive exact analytical solutions under arbitrary class imbalance. Contribution/Results: We systematically disentangle three distinct imbalance sources—intrinsic, training-set, and test-set—and prove that the optimal training imbalance ratio is generally not 50%, but instead evolves nontrivially with intrinsic imbalance, sample size, and noise level. We identify a sharp performance crossover between low-noise and high-noise regimes. Our theory quantitatively characterizes the individual impact of each imbalance source and yields actionable guidelines for constructing training sets—directly challenging the empirical “balance-is-optimal” heuristic. This work provides the first unified theoretical foundation for imbalanced learning in binary classification.

Technology Category

Application Category

📝 Abstract
Class imbalance (CI) is a longstanding problem in machine learning, slowing down training and reducing performances. Although empirical remedies exist, it is often unclear which ones work best and when, due to the lack of an overarching theory. We address a common case of imbalance, that of anomaly (or outlier) detection. We provide a theoretical framework to analyze, interpret and address CI. It is based on an exact solution of the teacher-student perceptron model, through replica theory. Within this framework, one can distinguish several sources of CI: either intrinsic, train or test imbalance. Our analysis reveals that the optimal train imbalance is generally different from 50%, with a non trivial dependence on the intrinsic imbalance, the abundance of data and on the noise in the learning. Moreover, there is a crossover between a small noise training regime where results are independent of the noise level to a high noise regime where performances quickly degrade with noise. Our results challenge some of the conventional wisdom on CI and offer practical guidelines to address it.
Problem

Research questions and friction points this paper is trying to address.

Imbalanced datasets
Machine learning
Anomaly detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Imbalanced Data
Anomaly Detection
Theoretical Framework
F
F. S. Pezzicoli
TAU team, LISN, Université Paris-Saclay, CNRS, Inria, 91405 Orsay, France
V
V. Ros
LPTMS, Université Paris Saclay, CNRS, 91405 Orsay, France; TAU team, LISN, Université Paris-Saclay, CNRS, Inria, 91405 Orsay, France
F
F. P. Landes
TAU team, LISN, Université Paris-Saclay, CNRS, Inria, 91405 Orsay, France
M
M. Baity-Jesi
SIAM Department, ETH Eawag, 8600 Dübendorf, Switzerland