Tightening Control in Neyman--Pearson Linear Classification

๐Ÿ“… 2026-07-03
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the challenge in Neymanโ€“Pearson (NP) classification of reliably constraining the type I error (i.e., misclassification rate of the prioritized class) under finite-sample settings, where existing methods often suffer from insufficient control. The authors identify this issue as stemming from an over-optimism bias inherent in statistical learning and propose two novel algorithms: one that enforces type I error control in expectation and another that guarantees such control with high probability. By integrating asymptotic theory, empirical risk minimization, and confidence bound construction, the proposed framework yields the first NP classifiers with rigorous theoretical guarantees and tight error control. It further incorporates class-specific accuracy estimation and uncertainty quantification. Experiments demonstrate that the methods substantially improve type I error control in finite samples and enhance screening reliability in real-world cancer detection tasks.
๐Ÿ“ Abstract
Neyman--Pearson classification prioritizes one class by constraining its accuracy above a prespecified level, and then takes the accuracy of the other class as the utility objective. This paradigm is well suited for disease screening and diagnosis, among other applications. Statistical learning under this framework is complicated since classifier performance determines its acceptability. Furthermore, no learned classifier that is consistent for the oracle classifier can guarantee satisfaction of the control constraint in finite samples. Classical learning theory targets a control-relaxed empirical utility maximization (EUM) classifier. However, even the EUM classifier fails to achieve the desired control level on average. We conjecture that this under-control phenomenon is a manifestation of the over-optimism bias well known in standard statistical learning, and develop asymptotic theory to confirm it. Motivated by this insight, we propose refined learning procedures under two accuracy control strategies for the prioritized class: one controlling accuracy in expectation and the other with high probability. We further develop training-data-based methods to predict and infer class-specific accuracies of the resulting classifiers. Simulation studies demonstrate favorable finite-sample performance, and we illustrate the proposed methods with an application to cancer detection.
Problem

Research questions and friction points this paper is trying to address.

Neyman-Pearson classification
accuracy control
finite-sample guarantee
class-specific accuracy
statistical learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Neyman-Pearson classification
over-optimism bias
accuracy control
finite-sample guarantee
class-specific accuracy inference
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
Y
Yijian Huang
Department of Biostatistics and Bioinformatics, Emory University, Atlanta, Georgia 30322, U.S.A.