🤖 AI Summary
This paper addresses the challenge of analytically constructing the optimal ROC curve in binary hypothesis testing when prior distributions are unknown or intractable. We propose the first maximum likelihood estimator for the ROC curve based on observed likelihood ratio samples (MLE-ROC). Unlike conventional approaches, MLE-ROC operates directly on likelihood ratio samples without requiring knowledge of the underlying data distributions and exhibits strong convergence under the Lévy metric. We establish theoretical consistency of its AUC estimator and demonstrate—via simulations—that it significantly outperforms the empirical ROC estimator, especially in small-sample and highly imbalanced settings with sparse negative instances, reducing estimation error by over 40%. The key contribution is the first formal parameterization of the ROC curve in the likelihood ratio domain within a maximum likelihood estimation framework, accompanied by rigorous asymptotic statistical guarantees.
📝 Abstract
The optimal receiver operating characteristic (ROC) curve, giving the maximum probability of detection as a function of the probability of false alarm, is a key information-theoretic indicator of the difficulty of a binary hypothesis testing problem (BHT). It is well known that the optimal ROC curve for a given BHT, corresponding to the likelihood ratio test, is theoretically determined by the probability distribution of the observed data under each of the two hypotheses. In some cases, these two distributions may be unknown or computationally intractable, but independent samples of the likelihood ratio can be observed. This raises the problem of estimating the optimal ROC for a BHT from such samples. The maximum likelihood estimator of the optimal ROC curve is derived, and it is shown to converge to the true optimal ROC curve in the Lévy metric, as the number of observations tends to infinity. A classical empirical estimator, based on estimating the two types of error probabilities from two separate sets of samples, is also considered. The maximum likelihood estimator is observed in simulation experiments to be considerably more accurate than the empirical estimator, especially when the number of samples obtained under one of the two hypotheses is small. The area under the maximum likelihood estimator is derived; it is a consistent estimator of the true area under the optimal ROC curve.