🤖 AI Summary
This study addresses the challenging problem of extreme binary classification where the false negative rate approaches zero. For the first time, this work formally defines the problem and proposes a novel paradigm with theoretical guarantees. Methodologically, an adaptive thresholding algorithm grounded in extreme value theory is designed to rigorously constrain false negative risk, while permutation testing is incorporated to enable highly reliable feature selection. Extensive experiments demonstrate that the proposed approach significantly outperforms existing state-of-the-art baselines across multiple real-world datasets. Furthermore, its application in cancer screening scenarios validates both its superior predictive performance and clinical interpretability.
📝 Abstract
While binary classification is one of the most extensively studied problems in machine learning,
the regime in which the goal is to learn a classifier with an almost zero false negative rate remains largely unexplored.
In this paper, we introduce the Extreme Binary Classification problem, where the objective is to learn a classifier whose false negative rate $α$ is constrained by $ε_{N_1}=o_{N_1\to\infty}(1/N_1)$, with $N_1$ denoting the number of positive examples in the training set.
To address this problem, we propose a threshold adaptation method theoretically grounded in guarantees derived from Extreme Value Theory, together with a feature selection procedure based on a permutation test applied to sample maxima.
Experimental results on four real-world datasets of varying sizes demonstrate that our approach compares favorably with state-of-the-art methods.
In addition, we illustrate its interpretability through an application to a cancer screening dataset.