🤖 AI Summary
This study addresses the challenge of optimizing partial Area Under the Curve (pAUC) in the absence of negative samples, as conventional approaches rely on fully labeled positive and negative data. To overcome this limitation, this work proposes a pAUC maximization framework that operates exclusively on Positive and Unlabeled (PU) data. By leveraging density ratio estimation, an unbiased risk estimator is derived, which is subsequently optimized via empirical risk minimization combined with smooth approximation techniques to handle the threshold-dependent pAUC objective function. The primary contribution lies in achieving, for the first time, the representation and estimation of threshold-based pAUC using solely PU data, thereby transcending the constraints of traditional dual-label paradigms. Extensive experiments conducted on ten real-world datasets thoroughly validate the effectiveness of the proposed method.
📝 Abstract
The partial area under the receiver operating characteristic curve (pAUC) is an important performance metric for binary classification that summarizes true positive rates within a specific range of false positive rates (FPRs). Classifiers that achieve high pAUC need to be obtained in many real-world applications such as cybersecurity, medical care, and advertising. Although many methods for maximizing the pAUC have been proposed, they typically require both labeled positive and negative data for training. However, in practice, labeled negative data are often difficult to collect due to privacy concerns or the need for high expertise to annotate them. In this paper, we propose a method for maximizing the pAUC from positive and unlabeled (PU) data without negative data. Within an empirical risk minimization framework, we show that the pAUC, including its FPR-dependent thresholds, can be represented using only the positive and marginal densities, and derive an empirical estimator from PU data. A classifier is then trained by maximizing the derived smoothed empirical pAUC estimator. We experimentally demonstrate the effectiveness of the proposed method with ten real-world datasets.