🤖 AI Summary
This work addresses the vulnerability of deep neural networks to label noise and the limited generalizability of existing data-cleaning methods that rely on handcrafted thresholds or single metrics. The authors propose an adaptive data-cleaning framework that, for the first time, integrates local KNN disagreement, global distance to cluster centroids, and normalized learning dynamics to construct a 2D/3D multi-dimensional feature representation. Within a unified low-dimensional space, robust sample selection is achieved via Gaussian mixture model clustering, eliminating the need to predefine noise ratios or thresholds. Evaluated on CIFAR-10, MNIST, and ImageNet-100, the method substantially improves cleaning recall—reaching over 98% on ImageNet-100 under 40% label noise—and effectively enhances downstream model performance.
📝 Abstract
Deep neural networks (DNNs) excel in computer vision tasks given large annotated datasets. In real-world applications, however, labels are often corrupted by ambiguity, human error, or dynamic environments. Over-parameterized DNNs easily memorize these noisy labels during training, degrading model accuracy and generalization. Existing data-cleaning and sample-selection strategies often rely on manually specified thresholds, prior knowledge of the noise ratio, or a single metric (either learning dynamics or geometric structure), making them unstable in complex data regimes. This paper proposes a self-adaptive data-cleaning framework that integrates local, global, and learning dynamics cues for robust noisy-label detection. Samples are mapped into a unified low-dimensional feature space through a modular feature concatenation paradigm. We provide two instantiations: a 2D metric integrating class-adaptive KNN-based local disagreement with k-means-based global centroid distance, and a 3D multi-metric that additionally incorporates a z-normalized score. Unlike conventional 1D Gaussian Mixture Models applied to a single scalar metric, our framework performs multi-metric clustering on the feature space to adaptively partition samples into clean-dominant and noise-dominant components without requiring manual thresholds or noise priors. Experiments on CIFAR-10, MNIST, and ImageNet-100 with 5% to 40% symmetric label noise show high recall across settings, including near-perfect recall (>=98%) on ImageNet-100 at 40% noise. Subsequent training yields accuracy gains across evaluated settings, especially under severe corruption on ImageNet-100. These findings suggest that multi-metric integration provides a threshold-free, practical, and low-tuning strategy for noisy label detection.