Score
Designs and evaluates methods that detect, filter, reweight, or replace incorrect labels in supervised datasets by estimating the most probable true label from the input features and observed noisy label, producing corrected training targets and confidence scores. Builds pipelines for relabeling, confident label cleaning, noisy-label filtering, or instance reweighting and analyzes how these corrections affect downstream model robustness and performance.
Real-world data often contain label noise that severely degrades model performance. This work proposes Relabeler, a novel framework that, for the first time, jointly models both local and global relationships among data instances within a unified architecture to enable end-to-end label correction. By leveraging input features and observed labels in concert, Relabeler estimates the most probable clean labels through a data-centric learning paradigm that integrates relational modeling, probabilistic label inference, and end-to-end optimization. Extensive experiments demonstrate that the method consistently outperforms existing approaches across diverse datasets, noise types, and noise rates, achieving up to a 58% improvement in label correction accuracy and a 6% gain in downstream task performance.
This work addresses a critical limitation in existing reconstruction-based methods for learning with noisy labels: their tendency to jointly assess the reliability of observed labels and pseudo-targets, which often leads to unreliable signals substituting one another and hinders effective denoising. To overcome this, the paper proposes TRACE, a novel framework that decouples the reliability evaluation of these two sources for the first time. Specifically, it evaluates observed labels through loss fitting, shallow-feature relational stability, and prediction consistency, while assessing pseudo-targets via model confidence. These independent reliability estimates are then used to separately govern label correction and sample reweighting. By preventing error propagation, TRACE generates more trustworthy pseudo-supervision, significantly outperforming current reconstruction-based approaches across multiple synthetic and real-world noisy benchmarks, and thereby enhancing model robustness and generalization.
This study addresses the low accuracy and poor robustness of label noise detection in machine learning. We propose a model calibration–based framework for identifying mislabeled samples. Methodologically, we systematically validate and leverage model calibration to enhance the reliability of predicted probabilities, designing an instance-level trust score generation mechanism that integrates multiple calibration strategies—including temperature scaling and vector scaling—to refine base model outputs. Our key contributions are: (1) establishing a positive correlation between model calibration quality and mislabeling detection performance; (2) significantly improving the stability of trust scores across diverse noise types and intensities; and (3) achieving an average 12.6% improvement in detection accuracy on multiple real-world datasets, thereby effectively supporting data cleaning and relabeling. The approach combines theoretical rigor with practical deployability in industrial settings.
To address performance degradation of deep models under label noise, this paper proposes a **selective noise correction method** that jointly detects and corrects noisy labels. First, it dynamically identifies potentially noisy samples based on loss distribution and separates clean and noisy subsets via a two-stage mechanism. Subsequently, it estimates a **local noise transition matrix** exclusively on the noisy subset, enabling end-to-end differentiable loss reweighting—thereby preserving all samples while avoiding overcorrection induced by global noise modeling. This is the first approach to achieve **dynamic selective correction**, balancing data integrity with correction precision. Extensive experiments on MNIST, CIFAR-10/100, and scRNA-seq cell-type annotation datasets demonstrate consistent improvements: average accuracy gains of 3.2–5.7% over state-of-the-art methods, alongside significantly enhanced robustness and generalization.
This work addresses the critical question of whether model retraining under randomly corrupted label noise can improve generalization accuracy. We propose a hard-label retraining mechanism and provide the first theoretical proof that, under the linearly separable data assumption, retraining with hard labels predicted by the model itself strictly improves generalization accuracy. Furthermore, we design a consensus-based retraining strategy that retains only samples for which the model’s predicted label agrees with the original (noisy) label—thereby enhancing both robustness and privacy preservation. Integrated within a label differential privacy (DP) training framework, this approach achieves over a 6% accuracy gain on CIFAR-100 using ResNet-18 under ε = 3 label DP, significantly improving learning efficacy under label noise and strengthening the privacy–utility trade-off.
Label noise in real-world datasets severely degrades model performance. To address this issue, this work proposes the CANOLA framework, which uniquely integrates explicit noise distribution modeling with progressive soft label correction. CANOLA employs a noise-aware deep neural network to estimate the underlying noise distribution and iteratively refines soft labels during training in a cautious manner, thereby avoiding premature or erroneous corrections that compromise stability and efficacy. Extensive experiments demonstrate that CANOLA significantly outperforms existing methods across six benchmark datasets, achieving relative error reductions of 19%–52%. Moreover, classifiers trained on labels refined by CANOLA surpass even sophisticated models by up to 67% in performance, highlighting the framework’s effectiveness in enhancing downstream learning tasks.
Label noise in supervised classification can severely degrade model performance, necessitating label-cleaning methods that require no prior knowledge of the noise characteristics. This work proposes an unsupervised noise identification framework that constructs data subsets via Bernoulli random sampling and leverages the linear relationship between cross-validation error and subset noise level to build a mixture distribution for distinguishing clean and noisy samples. Theoretically, we prove that under this sampling scheme, the mean noise level converges to two separable distributions, offering the first probabilistic coupling analysis guarantee for unsupervised label cleaning. The method is classifier-agnostic and consistently achieves significant improvements in both noise detection accuracy and downstream classification performance across synthetic and real-world datasets.
This study addresses the performance degradation of supervised learning on tabular data caused by latent label noise. To mitigate this issue, we propose a descriptor-driven adaptive ensemble framework. Specifically, the method constructs a data-centric inference module that integrates meta-models with confidence, neighborhood, and distribution detectors. By dynamically predicting the weights of these detectors, the framework enables automated data quality diagnosis and the selection of appropriate cleansing strategies. Experimental results demonstrate that the proposed approach achieves performance comparable to Confident Learning on standard benchmarks while significantly enhancing robustness across heterogeneous scenarios.
This work addresses the vulnerability of deep neural networks to label noise and the limited generalizability of existing data-cleaning methods that rely on handcrafted thresholds or single metrics. The authors propose an adaptive data-cleaning framework that, for the first time, integrates local KNN disagreement, global distance to cluster centroids, and normalized learning dynamics to construct a 2D/3D multi-dimensional feature representation. Within a unified low-dimensional space, robust sample selection is achieved via Gaussian mixture model clustering, eliminating the need to predefine noise ratios or thresholds. Evaluated on CIFAR-10, MNIST, and ImageNet-100, the method substantially improves cleaning recall—reaching over 98% on ImageNet-100 under 40% label noise—and effectively enhances downstream model performance.
This study investigates the impact of label noise on the generalization performance of language models, particularly in low-resource or high-noise settings. Focusing on Russian multi-domain text classification, it presents the first systematic comparison between Confident Learning and Dataset Cartography as automated methods for detecting label errors. Leveraging a fine-tuned rubert-base-cased model, the authors apply these techniques to filter noisy training data and validate their efficacy through controlled random-deletion baselines. Results demonstrate that Confident Learning substantially improves macro-F1 scores on small, high-noise datasets, whereas Dataset Cartography adopts a more conservative approach, removing fewer samples. Both methods consistently outperform random deletion, with their relative effectiveness closely dependent on dataset size and noise level.