🤖 AI Summary
This work addresses a critical yet overlooked issue in adaptive data cleaning: fluctuating sample removal counts caused by varying partition granularities introduce budget confounding bias, leading to spurious performance gains falsely attributed to contamination identification capability. To enable fair evaluation, the authors propose an operating-point-matched assessment framework that aligns removal budgets with recall rates and incorporates threshold-agnostic metrics (AUROC and AUPRC). They systematically uncover and resolve this budget confounding problem for the first time, introducing a multi-cue adaptive cleaner—integrating learning difficulty reweighting, Euclidean distance guidance, and fine-grained partitioning—and a false positive decomposition analysis. Experiments on CIFAR-10 and ImageNet-100 reveal that most existing methods lose their apparent advantage under matched operating points, demonstrating genuine efficacy only under low contamination rates or high-recall, heavily corrupted scenarios.
📝 Abstract
Adaptive data-cleaning methods replace manual filtering thresholds with data-driven partitions. However, changing the partition granularity, the number of groups used to segment samples by estimated corruption risk, can implicitly shift the decision boundary and alter the overall number of removed samples. This creates a bias known as removal-budget confounding, where apparent gains in metrics like precision or false-positive rate reflect a smaller removal budget rather than superior corruption discrimination. To address this evaluation bias, we introduce an operating-point-aware evaluation framework that evaluates methods using matched-budget and matched-recall controls alongside threshold-independent metrics (AUROC and AUPRC). We test this framework on a multi-cue adaptive cleaner redesign featuring a reweighted learning-difficulty cue, an auxiliary Euclidean-distance cue, and increased partition granularity intended to isolate clean-but-difficult samples. While naive evaluations (assessing configurations at their own induced operating points) suggest substantial performance improvements for the redesign, these gains disappear once operating points are equalized. False-positive decomposition reveals that clean-but-difficult samples primarily drive error counts at low corruption rates, become threshold-dependent at moderate corruption, and contribute negligibly under severe corruption. Experiments on CIFAR-10 and ImageNet-100 demonstrate that most performance differences observed in naive evaluation shrink or vanish at low-to-moderate corruption when operating points are matched. True ranking advantages only remain in specific low-prevalence settings and in high-recall regions under severe corruption. These findings highlight that adaptive cleaning methods must be benchmarked at matched operating points to ensure performance gains reflect genuine corruption discrimination.