🤖 AI Summary
Noisy labels in large-scale web data degrade model performance, while existing methods lack reliable criteria for sample selection. To address this, we propose a surrogate-model-driven robust sample selection framework. Our approach replaces conventional static thresholds—based on loss or confidence—with a lightweight, transferable surrogate model, enabling generalization across diverse network architectures and datasets. The framework jointly optimizes sample confidence estimation through surrogate distillation, consistency regularization, dynamic threshold calibration, and a noise-robust loss function, all in an end-to-end manner. Extensive experiments on standard noisy-label benchmarks—including CIFAR-10, CIFAR-100, and WebVision—demonstrate that our method consistently outperforms state-of-the-art approaches such as FixMatch and Co-teaching, achieving absolute accuracy improvements of 3.2%–5.8%.