auroc and auprc evaluation

Designs and implements evaluation pipelines and analyses that compute and visualize receiver operating characteristic (ROC) and precision–recall curves, associated summary metrics (AUROC/AUPRC/AUC), equal error rates, and threshold trade-off statistics to quantify classifier or anomaly-detection performance across decision thresholds. Performs statistical estimation and comparison (confidence intervals, significance tests), measures incremental AUC gains from added features or markers, builds cross-dataset and incremental assessment protocols, and accounts for class imbalance and inter-feature correlations when comparing models or reporting performance.

aurocandauprcevaluation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$199K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

A Closer Look at AUROC and AUPRC under Class Imbalance

Jan 11, 2024
MB
Matthew B. A. McDermott
🏛️ Harvard | Aarhus University | Massachusetts Institute of Technology | IRCCS Humanitas Research Hospital

This paper challenges the widely held assumption in machine learning that the Area Under the Precision-Recall Curve (AUPRC) is universally superior to the Area Under the ROC Curve (AUROC) under class imbalance—and investigates its implications for algorithmic fairness. Method: We conduct rigorous theoretical analysis, experiments on semi-synthetic and real-world fairness-sensitive datasets, and a large-scale bibliometric study covering over one million publications. Contribution/Results: We provide the first formal proof that AUPRC is not generally advantageous under extreme imbalance; instead, it systematically amplifies group-level bias by favoring subpopulations with higher positive-class density. We trace and empirically refute the long-standing misconception that “AUPRC is inherently better.” Furthermore, we establish verifiable criteria delineating the applicability boundaries of AUROC versus AUPRC, and propose a principled, imbalance- and subgroup-aware framework for metric selection—thereby offering both theoretical foundations and practical guidance for fair model evaluation.

Algorithmic FairnessAUPRC vs AUROCImbalanced Classes

Decomposing Global AUC into Cluster-Level Contributions for Localized Model Diagnostics

Aug 10, 2025
AS
Agus Sudjianto
🏛️ H2O.ai | 2nd Order Solutions

Global AUC fails to reveal localized model performance deficiencies within subpopulations. To address this, we propose the first decomposable cluster-level AUC evaluation framework, which rigorously decomposes global AUC into two orthogonal components: intra-cluster ranking ability and inter-cluster discriminative ability. Methodologically, leveraging the geometric properties of ROC curves and cluster structure, we derive an additive, unbiased AUC decomposition formula and provide theoretical guarantees for its interpretability and statistical consistency. Unlike conventional metrics such as Brier score or log loss, our framework is the first to enable AUC decomposition at the cluster granularity, facilitating fine-grained diagnostic analysis and high-risk subgroup identification. Empirical evaluations on credit approval and fraud detection tasks demonstrate significant improvements in model validation accuracy and risk management efficacy.

Compare AUC with decomposable metrics like Brier scoreDecompose global AUC for localized model diagnosticsEvaluate classifier performance within and across clusters

To address the critical “zero false negatives” (i.e., zero missed detections) requirement in high-stakes domains such as industrial monitoring, healthcare, and cybersecurity, this paper proposes a trustworthy training paradigm. We introduce a differentiable approximate partial AUC loss (tapAUC), which concentrates on the low false positive rate (FPR) region of the ROC curve to maximize true positive rate (TPR) under a strict zero-FN constraint, while adaptively determining a robust decision threshold. Evaluated on six benchmark datasets, our method achieves an average TPR of 92.52% at FPR = 20.43%, outperforming state-of-the-art methods by 4.3 percentage points. It is the first approach to significantly enhance detection sensitivity while provably guaranteeing zero missed anomalies—thereby establishing a new anomaly detection framework that combines theoretical soundness with practical applicability for safety-critical applications.

Improve True Positive Rate with controlled False Positive RateMinimize false negatives using partial AUC ROCOptimize anomaly detection for critical applications

This study addresses the overreliance on the Area Under the ROC Curve (AUC) in software defect prediction research, which can lead to biased model evaluation as AUC fails to capture a model’s discriminative performance across all classification thresholds. To overcome this limitation, the authors propose an augmented ROC curve and alternative visualization techniques that explicitly model the true positive rate and false positive rate as functions of the decision threshold, annotating corresponding threshold points directly on the ROC curve. Their analysis demonstrates that a high AUC does not guarantee superior performance over random guessing at every threshold, thereby exposing critical shortcomings of conventional evaluation practices. The work underscores the necessity of incorporating multi-threshold classification performance into a more comprehensive and nuanced assessment framework.

AUCModel EvaluationROC Curve

Preserving AUC Fairness in Learning with Noisy Protected Groups

May 24, 2025
MW
Mingyang Wu
🏛️ Purdue University | Florida International University | University at Albany | State University of New York | Amazon

When protected attributes (e.g., gender, race) are corrupted by label noise, AUC-based fairness metrics—particularly the AUC difference (ΔAUC)—are highly susceptible to degradation. Existing methods assume clean sensitive attributes, limiting their practical robustness. Method: This paper proposes the first theoretically grounded robust AUC fairness optimization framework. Departing from prior assumptions, it introduces distributionally robust optimization (DRO) into AUC fairness learning, deriving a provable upper bound on ΔAUC under attribute noise. The method integrates AUC gradient approximation, noise-robust loss design, and multi-task fairness constraints to ensure robust modeling despite noisy protected attributes. Results: Extensive experiments on tabular and image datasets demonstrate that our approach significantly outperforms state-of-the-art methods, reducing average ΔAUC by 37% while preserving baseline AUC performance—confirming both fairness improvement and predictive utility.

Addressing biases in AUC optimization for protected groupsEnsuring AUC fairness with noisy protected groupsRobust AUC fairness under noisy group labels

Latest Papers

What's happening recently
View more

This study addresses the challenge posed by severe class imbalance in anomaly detection, which complicates the interpretation and comparison of common evaluation metrics. The authors systematically analyze the behavior of AUROC, AUPR, F1-score, and Matthews Correlation Coefficient (MCC) across varying anomaly ratios and introduce a novel "metric landscape" visualization technique. This approach reveals, for the first time, each metric’s inherent preference for true positive rate versus true negative rate and how their stability varies with imbalance levels. By modeling the relationship between metrics and anomaly prevalence, the work delineates clear applicability boundaries for each metric, thereby providing a principled, interpretable foundation for reliable metric selection in highly imbalanced anomaly detection scenarios.

anomaly detectionclass imbalanceevaluation metrics

This study addresses the critical limitation of current video anomaly detection methods, which perform adequately within the same scene but suffer severe performance degradation when deployed across different scenes—a shortcoming obscured by the prevailing evaluation practice of reporting in-dataset AUC. The authors conduct the first systematic audit and quantification of this cross-domain performance collapse, proposing an unsupervised normality model built upon frozen off-the-shelf visual features (e.g., CLIP, DINOv2, ResNet-50, EfficientNet-B0) combined with nearest-neighbor search and PaDiM-style Mahalanobis distance for anomaly scoring. Experiments reveal that while intra-scene average AUC reaches 0.704, it plummets to 0.499—near random—in cross-scene settings, with DINOv2, the best-performing backbone, exhibiting the most pronounced drop and yielding a practical false alarm rate as high as 31,931 alerts per hour, thereby challenging the validity of existing evaluation paradigms.

AUCcross-dataset generalizationout-of-distribution

This study addresses label circularity bias in multi-label chest X-ray classification caused by the conflation of prospective clinical indications with retrospective radiology report sections (Findings/Impression). The authors propose a rigorous evaluation framework that strictly separates these two sources of information. Using patient-level clustering bootstrapping and frozen DenseNet-121 and Bio+ClinicalBERT encoders combined with CheXpert labels, they systematically compare multiple fusion strategies. They introduce SectionGuard-MI, a novel permutation-aware method that significantly outperforms baselines in AUPRC (+0.0289, p=0.004). DeepSets achieves the best AUROC (0.787) under purely prospective settings, while leveraging full reports yields an inflated AUROC of 0.979—quantifying, for the first time, the extent of information leakage and establishing the true predictive value of prospective clinical indications.

chest radiograph classificationclinical indicationmultimodal fusion

This study addresses the critical gap that clinical machine learning models, despite achieving high AUROC scores (0.849–0.900), often produce predictions inconsistent with medical common sense and lack methods for verifying the reasonableness of individual predictions without ground-truth labels. The authors propose the first systematic framework of 12 metamorphic relations (MRs) derived from authoritative clinical guidelines for ICU prediction tasks, accompanied by a five-tier validation strategy to ensure their clinical validity. Evaluations on MIMIC-III/IV and UCI Heart Disease datasets, using metamorphic testing and fault injection, reveal that MT violation rates range from 27% to 87%. Notably, injecting sign errors into blood pressure features leaves AUROC nearly unchanged but increases MT violation rates by 31–67 percentage points, effectively exposing behavioral flaws invisible to conventional performance metrics.

behavioral correctnessclinical machine learningmedical knowledge consistency

Hot Scholars

DT

Daniel Truhn

Professor of Radiology, University Hospital Aachen
Machine LearningArtificial IntelligenceComputer VisionMedical Imaging
SN

Sven Nebelung

Department of Diagnostic and Interventional Radiology, University Hospital Aachen
Advanced MRI TechniquesFunctionality AssessmentBiomechanical ImagingCartilage
SH

Shenda Hong

Assistant Professor, Peking University
AI ECGBiosignalAI for Digital HealthHealth Data Science
KZ

Ke Zou

Apple, Inc
Power electronicsSwitched-capacitor ConverterPower Semiconductor Devices