Score
Designs and implements methods and tools that identify, flag, and rank dataset examples whose annotated labels are likely incorrect or corrupted, producing per-sample noisiness scores or prioritized lists for expert review. These methods analyze local instance relationships and global dataset patterns to estimate label confidence, support filtering or relabeling workflows, and reduce label noise to improve downstream model training.
Widespread label noise (10–25%) in NLP benchmark datasets leads to systematic underestimation of model performance, with many purported “LLM failures” attributable to annotation errors rather than model limitations. Method: We propose LLM-as-a-judge—a framework leveraging ensemble judgments from GPT-4, Claude, and Llama, combined with consistency voting and error-sensitivity analysis to automatically detect mislabeled instances; we further apply label smoothing and confident learning for robust label recalibration. Contribution/Results: Comprehensive evaluation across the TRUE benchmark suite reveals substantial disparities in quality and efficiency among expert, crowdsourced, and LLM-generated annotations. After correction, state-of-the-art models achieve average accuracy gains of 3.2–7.8 percentage points. This work provides the first empirical evidence of systematic label-noise interference in LLM evaluation and introduces a scalable, collaborative adjudication paradigm that reframes data correction as model performance recalibration.
This paper addresses binary classification under noisy labels by proposing a robust learning method based on hypergraph normalized cut (HNC). The core contribution is the first integration of a learnable confidence-weighting mechanism into the HNC framework, enabling the model to adaptively identify and rectify erroneous labels without requiring prior knowledge of the noise rate. The method achieves efficient joint optimization of classification and noise detection through a parameterized network-flow formulation of the minimum cut. Extensive experiments on both synthetic and real-world noisy-label benchmarks demonstrate that the proposed approach significantly improves classification accuracy and exhibits strong capability in identifying corrupted samples, consistently outperforming state-of-the-art robust classification methods.
In noisy label learning, sample selection suffers from dual biases: data bias (imbalanced selection sets) and training bias (error accumulation). To address these issues, this paper proposes ITEM, a noise-tolerant expert model. ITEM is the first to jointly model and mitigate both biases in a unified framework. It introduces a lightweight multi-expert robust network architecture, integrated with a dual-weighted class-discriminative sampler and a hybrid mini-batch training strategy. Additionally, an error-robust optimization mechanism is incorporated to enhance generalization under label noise. Extensive experiments on multiple benchmark datasets with synthetic and real-world label noise demonstrate that ITEM consistently outperforms state-of-the-art methods, achieving average accuracy gains of 3–5% while reducing parameter count by over 20%. The source code is publicly available.
This work investigates the robustness of Gradient Boosting Decision Trees (GBDTs) to label noise in tabular classification. Addressing the sensitivity of conventional GBDTs to noisy labels and their lack of intrinsic noise detection mechanisms, we propose Gradients—a novel gradient-based noise detection method inspired by deep learning paradigms—and design a GBDT-specific noise-handling pipeline incorporating dynamic relabeling and early stopping. Implemented on XGBoost and LightGBM, our approach jointly leverages prediction confidence, sample consistency, and gradient-derived features for noise identification. Experiments demonstrate state-of-the-art performance: Gradients achieves 99.1% noise detection accuracy on the Adult dataset, substantially outperforming baseline methods. Furthermore, robustness and generalizability are validated across multiple benchmark datasets—including Covertype and Breast Cancer—under diverse noise settings. To our knowledge, this is the first work to adapt deep-learning-inspired noise detection principles to GBDT frameworks, establishing a principled and effective approach for enhancing GBDT reliability in real-world, noisy tabular learning scenarios.
This study addresses the fundamental tension between strong aggregate performance and frequent individual-level misclassifications in machine learning models under label noise. We introduce “individual-level regret”—a novel metric quantifying unforeseen misclassifications attributable to label corruption. To mitigate this issue, we propose a robust modeling framework grounded in denoised dataset resampling, integrating ensemble learning, counterfactual data generation, robust empirical risk minimization, and uncertainty calibration—thereby enabling estimable individual error probabilities. Evaluated across multiple clinical prediction tasks, our approach significantly reduces volatility in individual-level errors, enhancing model reliability and clinical deployability. Empirical results demonstrate a 37–62% reduction in regretful misclassifications compared to baseline methods.
This study investigates the impact of label noise on the generalization performance of language models, particularly in low-resource or high-noise settings. Focusing on Russian multi-domain text classification, it presents the first systematic comparison between Confident Learning and Dataset Cartography as automated methods for detecting label errors. Leveraging a fine-tuned rubert-base-cased model, the authors apply these techniques to filter noisy training data and validate their efficacy through controlled random-deletion baselines. Results demonstrate that Confident Learning substantially improves macro-F1 scores on small, high-noise datasets, whereas Dataset Cartography adopts a more conservative approach, removing fewer samples. Both methods consistently outperform random deletion, with their relative effectiveness closely dependent on dataset size and noise level.
Real-world data often contain label noise that severely degrades model performance. This work proposes Relabeler, a novel framework that, for the first time, jointly models both local and global relationships among data instances within a unified architecture to enable end-to-end label correction. By leveraging input features and observed labels in concert, Relabeler estimates the most probable clean labels through a data-centric learning paradigm that integrates relational modeling, probabilistic label inference, and end-to-end optimization. Extensive experiments demonstrate that the method consistently outperforms existing approaches across diverse datasets, noise types, and noise rates, achieving up to a 58% improvement in label correction accuracy and a 6% gain in downstream task performance.
This work addresses a critical limitation in existing reconstruction-based methods for learning with noisy labels: their tendency to jointly assess the reliability of observed labels and pseudo-targets, which often leads to unreliable signals substituting one another and hinders effective denoising. To overcome this, the paper proposes TRACE, a novel framework that decouples the reliability evaluation of these two sources for the first time. Specifically, it evaluates observed labels through loss fitting, shallow-feature relational stability, and prediction consistency, while assessing pseudo-targets via model confidence. These independent reliability estimates are then used to separately govern label correction and sample reweighting. By preventing error propagation, TRACE generates more trustworthy pseudo-supervision, significantly outperforming current reconstruction-based approaches across multiple synthetic and real-world noisy benchmarks, and thereby enhancing model robustness and generalization.
Label noise in real-world datasets severely degrades model performance. To address this issue, this work proposes the CANOLA framework, which uniquely integrates explicit noise distribution modeling with progressive soft label correction. CANOLA employs a noise-aware deep neural network to estimate the underlying noise distribution and iteratively refines soft labels during training in a cautious manner, thereby avoiding premature or erroneous corrections that compromise stability and efficacy. Extensive experiments demonstrate that CANOLA significantly outperforms existing methods across six benchmark datasets, achieving relative error reductions of 19%–52%. Moreover, classifiers trained on labels refined by CANOLA surpass even sophisticated models by up to 67% in performance, highlighting the framework’s effectiveness in enhancing downstream learning tasks.
This work addresses the high cost of machine learning benchmarking by proposing a systematic framework to efficiently select small, representative subsets of datasets while preserving model ranking stability. The study presents the first comprehensive evaluation of various dataset selection strategies—including clustering, A/D-optimal experimental designs, random baselines, and a greedy farthest-first (FAFI) approach—on rank fidelity. It derives a theoretical upper bound on Spearman rank correlation error for FAFI and integrates bootstrap aggregation to yield statistically rigorous confidence intervals for comparing strategy performance. Empirical results demonstrate that as few as five datasets suffice to achieve 0.95 rank correlation in time series classification, significantly outperforming random selection in NLP tasks, though gains are limited in recommendation systems.