Score
Designs and implements methods that estimate per-annotator error patterns (e.g., confusion matrices) and systematic biases, produce debiased aggregated labels from noisy or “silver” annotations, and incorporate historical gold data or partial overlap among annotators to improve label calibration.
This study addresses the low accuracy and poor robustness of label noise detection in machine learning. We propose a model calibration–based framework for identifying mislabeled samples. Methodologically, we systematically validate and leverage model calibration to enhance the reliability of predicted probabilities, designing an instance-level trust score generation mechanism that integrates multiple calibration strategies—including temperature scaling and vector scaling—to refine base model outputs. Our key contributions are: (1) establishing a positive correlation between model calibration quality and mislabeling detection performance; (2) significantly improving the stability of trust scores across diverse noise types and intensities; and (3) achieving an average 12.6% improvement in detection accuracy on multiple real-world datasets, thereby effectively supporting data cleaning and relabeling. The approach combines theoretical rigor with practical deployability in industrial settings.
This work investigates the intrinsic mechanism enabling deep models to generalize well under label noise. Theoretically, we show that label noise primarily perturbs low-order singular components of weight matrices, while the dominant subspace governing generalization—the principal subspace spanned by top singular vectors—remains stably aligned. This yields the first rigorous subspace-level characterization of generalization robustness under label corruption. Building on this insight, we propose LIP (Low-rank Invariant Projection), a lightweight plug-in module that explicitly regularizes the principal subspace structure of weight matrices during training. Extensive experiments across diverse synthetic and real-world noisy-label benchmarks demonstrate that LIP consistently improves classification accuracy of mainstream architectures—including ResNet and ViT—without architectural modification. Our results empirically and theoretically establish principal subspace stability as the fundamental geometric principle underlying label-noise robustness.
In noisy label learning, sample selection suffers from dual biases: data bias (imbalanced selection sets) and training bias (error accumulation). To address these issues, this paper proposes ITEM, a noise-tolerant expert model. ITEM is the first to jointly model and mitigate both biases in a unified framework. It introduces a lightweight multi-expert robust network architecture, integrated with a dual-weighted class-discriminative sampler and a hybrid mini-batch training strategy. Additionally, an error-robust optimization mechanism is incorporated to enhance generalization under label noise. Extensive experiments on multiple benchmark datasets with synthetic and real-world label noise demonstrate that ITEM consistently outperforms state-of-the-art methods, achieving average accuracy gains of 3–5% while reducing parameter count by over 20%. The source code is publicly available.
Widespread label noise (10–25%) in NLP benchmark datasets leads to systematic underestimation of model performance, with many purported “LLM failures” attributable to annotation errors rather than model limitations. Method: We propose LLM-as-a-judge—a framework leveraging ensemble judgments from GPT-4, Claude, and Llama, combined with consistency voting and error-sensitivity analysis to automatically detect mislabeled instances; we further apply label smoothing and confident learning for robust label recalibration. Contribution/Results: Comprehensive evaluation across the TRUE benchmark suite reveals substantial disparities in quality and efficiency among expert, crowdsourced, and LLM-generated annotations. After correction, state-of-the-art models achieve average accuracy gains of 3.2–7.8 percentage points. This work provides the first empirical evidence of systematic label-noise interference in LLM evaluation and introduces a scalable, collaborative adjudication paradigm that reframes data correction as model performance recalibration.
Label noise significantly degrades model performance, necessitating effective detection methods. This work proposes Adaptive Label Error Detection (ALED), which first extracts and denoises intermediate features from deep convolutional networks, then models each class as a multivariate Gaussian distribution on a low-dimensional manifold, and finally identifies mislabeled samples via a Bayesian likelihood ratio test. ALED is the first approach to integrate feature denoising, class-conditional Gaussian modeling, and likelihood ratio testing within a unified framework. Evaluated on multiple medical imaging datasets, it substantially outperforms existing techniques; fine-tuning models with labels corrected by ALED reduces test error rates by 33.8%, achieving both high sensitivity and precision.
This work addresses a critical limitation in existing reconstruction-based methods for learning with noisy labels: their tendency to jointly assess the reliability of observed labels and pseudo-targets, which often leads to unreliable signals substituting one another and hinders effective denoising. To overcome this, the paper proposes TRACE, a novel framework that decouples the reliability evaluation of these two sources for the first time. Specifically, it evaluates observed labels through loss fitting, shallow-feature relational stability, and prediction consistency, while assessing pseudo-targets via model confidence. These independent reliability estimates are then used to separately govern label correction and sample reweighting. By preventing error propagation, TRACE generates more trustworthy pseudo-supervision, significantly outperforming current reconstruction-based approaches across multiple synthetic and real-world noisy benchmarks, and thereby enhancing model robustness and generalization.
This study addresses the bias in classifier training and statistical inference caused by noisy human-reviewed labels. To mitigate this issue, the authors propose Partially Adjudicated Design-based Supervised Learning (PA-DSL), a novel framework that integrates partial expert adjudication into design-based supervised learning. By combining probability sampling audits, label noise correction, and design-weighted estimation, PA-DSL leverages recoverable signals from noisy labels while ensuring unbiased estimation. The method is applicable to various downstream tasks where audit and adjudication probabilities are known. Experiments on synthetic data and semi-synthetic Wikipedia Detox datasets demonstrate that, compared to approaches using only adjudicated labels, PA-DSL reduces root mean squared error by 10%–17% while maintaining nominal coverage.
This work addresses the challenge of evaluating generative models under limited access to expensive expert annotations (gold labels), where widely used crowd-sourced or vendor-provided labels (silver labels) are often noisy. Direct aggregation of such silver labels introduces bias, while estimators relying on sparse gold labels suffer from high variance, hindering accurate model comparison. To overcome these limitations, the paper proposes HERO, a novel framework that systematically leverages historical evaluation data to calibrate the reliability of current-round silver annotators. By anchoring high-fidelity covariate information from past evaluations, HERO constructs a robust estimator with both low bias and low variance. The method remains effective even when only a subset of historical annotators participate in new evaluation rounds, and theoretical analysis alongside experiments on synthetic and real-world benchmarks demonstrates its significant improvement in assessment accuracy and stability.
Label noise in supervised classification can severely degrade model performance, necessitating label-cleaning methods that require no prior knowledge of the noise characteristics. This work proposes an unsupervised noise identification framework that constructs data subsets via Bernoulli random sampling and leverages the linear relationship between cross-validation error and subset noise level to build a mixture distribution for distinguishing clean and noisy samples. Theoretically, we prove that under this sampling scheme, the mean noise level converges to two separable distributions, offering the first probabilistic coupling analysis guarantee for unsupervised label cleaning. The method is classifier-agnostic and consistently achieves significant improvements in both noise detection accuracy and downstream classification performance across synthetic and real-world datasets.