🤖 AI Summary
The area under the ROC curve (AUC) is commonly interpreted as the probability that a classifier ranks a randomly chosen positive instance higher than a negative one; however, this interpretation relies on specific assumptions whose violation can introduce bias that has not been rigorously characterized. This work systematically reviews the relevant literature and, drawing on probabilistic and statistical methods, provides the first rigorous proof of the conditions under which this probabilistic interpretation holds. Furthermore, when these assumptions are violated, the study derives a computable upper bound on the resulting deviation. By establishing a solid theoretical foundation for the probabilistic interpretation of AUC and offering explicit error bounds, this research significantly enhances the reliability and practical applicability of ROC analysis.
📝 Abstract
The Receiver Operating Characteristic (ROC) curve of a binary classifier has often been utilized to measure the performance of the classifier. The area beneath this curve is used in particular because of its quoted probabilistic interpretation as being equal to the probability that the classifier will rank a random positive observation above a random negative observation. This paper formalizes this claim, produces a bound on how far away from the truth it is if a hypothesis is not met, and gives a small literature review of the ROC curve.