label noise characterization

Designs and implements analyses and tools to detect, localize, quantify, and model errors in dataset labels and annotations. This includes estimating per-instance and per-class noise rates and confusion patterns, fitting label-transition matrices, and producing dataset-specific noise profiles and annotation-error maps for both categorical and structured labels (e.g., segmentation masks).

labelnoisecharacterization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$207K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Learning to Detect Label Errors by Making Them: A Method for Segmentation and Object Detection Datasets

Aug 25, 2025
SP
Sarina Penquitt
🏛️ University of Wuppertal | Technical University of Berlin | Karlsruhe Institute of Technology (KIT) | Osnabrück University

Existing label-noise detection methods suffer from task specificity, limited learnability, and poor generalization across vision tasks. To address this, we propose the first unified, learnable framework for label-noise detection applicable to object detection, semantic segmentation, and instance segmentation. Our core innovation is an “error-correcting-via-error” paradigm: we reformulate label-noise detection as an instance segmentation problem by injecting controllable synthetic errors, and introduce a composite input representation alongside multi-task joint training. We establish a new authoritative benchmark on Cityscapes containing 459 real-world mislabeled instances. Extensive experiments demonstrate that our method significantly outperforms prior approaches, exhibiting strong generalization and robustness in both synthetic and real-world settings. This work provides a reproducible, scalable paradigm and empirical foundation for improving data quality in supervised learning.

Detecting label errors in object detection and segmentation datasetsImproving dataset quality through learning-based error injectionUnifying error detection across multiple computer vision tasks

From Label Error Detection to Correction: A Modular Framework and Benchmark for Object Detection Datasets

Aug 06, 2025
SP
Sarina Penquitt
🏛️ University of Wuppertal | Quality Match | Osnabrück University

This work addresses pervasive labeling errors—such as missing annotations, misclassifications, and imprecise bounding boxes—in object detection datasets. We propose REC✓D, a semi-automated correction framework that leverages pre-trained detectors to generate candidate mislabeling suggestions, employs lightweight crowdsourced micro-tasks for independent human verification by multiple annotators, and applies response aggregation to quantify annotation ambiguity and enhance correction robustness. To our knowledge, REC✓D is the first scalable, end-to-end system for detecting and correcting labels in object detection data. As a key contribution, we release a rigorously validated high-quality subset of pedestrian annotations from KITTI as a new benchmark. Experiments demonstrate that existing methods fail to detect up to 66% of ground-truth labeling errors, whereas REC✓D identifies and rectifies at least 24% of original annotation errors at a fraction of the cost of full manual re-annotation, substantially improving dataset quality and model evaluation reliability.

Address missing and inaccurate annotations in benchmark datasetsDetect and correct label errors in object detection datasetsImprove label quality through crowdsourced verification

Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performance

Oct 24, 2024
ON
Omer Nahum
🏛️ Technion - Institute of Technology | Google Research

Widespread label noise (10–25%) in NLP benchmark datasets leads to systematic underestimation of model performance, with many purported “LLM failures” attributable to annotation errors rather than model limitations. Method: We propose LLM-as-a-judge—a framework leveraging ensemble judgments from GPT-4, Claude, and Llama, combined with consistency voting and error-sensitivity analysis to automatically detect mislabeled instances; we further apply label smoothing and confident learning for robust label recalibration. Contribution/Results: Comprehensive evaluation across the TRUE benchmark suite reveals substantial disparities in quality and efficiency among expert, crowdsourced, and LLM-generated annotations. After correction, state-of-the-art models achieve average accuracy gains of 3.2–7.8 percentage points. This work provides the first empirical evidence of systematic label-noise interference in LLM evaluation and introduces a scalable, collaborative adjudication paradigm that reframes data correction as model performance recalibration.

Comparing annotation quality from experts, crowdsourcing, and LLMsDetecting label errors in NLP benchmark datasetsMitigating mislabeled data effects on model performance

This study addresses the pervasive issues of semantic mislabeling and bounding box localization errors in object detection datasets by systematically evaluating, for the first time, the effectiveness of training-free feature-space methods for annotation error detection. Leveraging multiple pretrained embedding models, the approach is rigorously tested on both synthetic noise—including symmetric, asymmetric, and localization-type perturbations—and real-world annotation errors in the VOC2012 and KITTI datasets. Experimental results demonstrate that feature-space methods are highly effective at identifying semantic mislabels but exhibit limited capability in detecting localization inaccuracies. To facilitate future research, the authors publicly release all code and a curated set of verified erroneous annotations, establishing a valuable benchmark for the community.

annotation errorsfeature-space methodsobject detection

To address the limitation of prior-dependent constraints undermining generalizability in hierarchical multi-label classification (HMC), this paper proposes EDR, a prior-free error-driven constraint discovery framework. EDR automatically identifies model misprediction patterns to induce interpretable, structured logical constraints; it further integrates a constraint-driven post-processing mechanism with a neuro-symbolic joint modeling architecture to jointly perform error detection, constraint recovery, and multi-level consistency verification. For the first time, EDR achieves fully automated, interpretable knowledge discovery and robust cross-domain constraint transfer without predefined constraints—enabling effective constraint learning even under label noise. Evaluated on multiple public benchmarks and a newly constructed military vehicle recognition dataset, EDR achieves an error detection F1-score exceeding 0.89 and constraint recovery accuracy above 92%, significantly improving both hierarchical consistency and overall classification performance.

Detects errors in hierarchical multi-label classification without prior constraintsLearns explainable rules for machine learning model failure modesRecovers constraints for neurosymbolic models from error detection rules

Latest Papers

What's happening recently
View more

This study investigates the impact of label noise on the generalization performance of language models, particularly in low-resource or high-noise settings. Focusing on Russian multi-domain text classification, it presents the first systematic comparison between Confident Learning and Dataset Cartography as automated methods for detecting label errors. Leveraging a fine-tuned rubert-base-cased model, the authors apply these techniques to filter noisy training data and validate their efficacy through controlled random-deletion baselines. Results demonstrate that Confident Learning substantially improves macro-F1 scores on small, high-noise datasets, whereas Dataset Cartography adopts a more conservative approach, removing fewer samples. Both methods consistently outperform random deletion, with their relative effectiveness closely dependent on dataset size and noise level.

data qualitylabel errorslanguage models

Real-world data often contain label noise that severely degrades model performance. This work proposes Relabeler, a novel framework that, for the first time, jointly models both local and global relationships among data instances within a unified architecture to enable end-to-end label correction. By leveraging input features and observed labels in concert, Relabeler estimates the most probable clean labels through a data-centric learning paradigm that integrates relational modeling, probabilistic label inference, and end-to-end optimization. Extensive experiments demonstrate that the method consistently outperforms existing approaches across diverse datasets, noise types, and noise rates, achieving up to a 58% improvement in label correction accuracy and a 6% gain in downstream task performance.

data qualitylabel correctionlabel corruption

Label noise in supervised classification can severely degrade model performance, necessitating label-cleaning methods that require no prior knowledge of the noise characteristics. This work proposes an unsupervised noise identification framework that constructs data subsets via Bernoulli random sampling and leverages the linear relationship between cross-validation error and subset noise level to build a mixture distribution for distinguishing clean and noisy samples. Theoretically, we prove that under this sampling scheme, the mean noise level converges to two separable distributions, offering the first probabilistic coupling analysis guarantee for unsupervised label cleaning. The method is classifier-agnostic and consistently achieves significant improvements in both noise detection accuracy and downstream classification performance across synthetic and real-world datasets.

Bernoulli samplingclassification errorlabel noise

This work addresses the gap between theory and practice in error detection and cleaning for tabular data by proposing and implementing an interactive web-based demonstration system that integrates error modeling, injection, and intelligent cleaning. For the first time, the system unifies machine learning–based data cleaning methods with dependency-aware error generation models within a visual platform, enabling users to upload tables, inject realistic errors, and interactively experience the cleaning process alongside interpretable explanations of the underlying mechanisms. By synergizing machine learning, database technologies, and web development, the project enhances intuitive understanding of error repair principles and provides a publicly accessible demo platform (https://cured.demo.calgo-lab.de/) that serves as an effective validation tool for both research and practical applications in data cleaning.

data cleaningerror detectionerror models

This study addresses the pervasive issue of data errors in real-world databases—such as missing values, redundancy, statistical biases, and outliers—which significantly degrade downstream analytical and machine learning performance. Recognizing that existing taxonomies are incomplete and terminology inconsistent, this work presents the first unified framework that integrates traditional data errors with statistically oriented inaccuracies critical in the AI era. It proposes a non-overlapping tripartite classification structure—comprising missing, erroneous, and redundant data—and systematically constructs a comprehensive catalog of 35 distinct error types. Through formal definitions, illustrative examples, and a thorough literature review, the paper establishes standardized terminology and precise characterizations, thereby offering a clear, rigorous theoretical foundation and practical toolkit for data quality assessment and cleaning.

data errorsdata qualityerror taxonomy

Hot Scholars

MS

Masashi Sugiyama

Director, RIKEN Center for Advanced Intelligence Project / Professor, The University of Tokyo
Machine LearningData MiningArtificial Intelligence
KY

Kailun Yang

Professor. School of Artificial Intelligence and Robotics, Hunan University (HNU); KIT; UAH; ZJU
Computer VisionComputational OpticsIntelligent VehiclesAutonomous Driving
KP

Kunyu Peng

Karlsruhe Institute of Technology
video understandingopen set recognitiongeneralizable deep learning
JZ

Junwei Zheng

CV:HCI, KIT; CVG, ETH Zurich
Visual LocalizationScene UnderstandingAssistive Technology
DW

Di Wen

Karlsruhe Institute of Technology
Fine-grained Action UnderstandingAnomaly DetectionRobustnessUncertainty