Score
Designs and implements analyses and tools to detect, localize, quantify, and model errors in dataset labels and annotations. This includes estimating per-instance and per-class noise rates and confusion patterns, fitting label-transition matrices, and producing dataset-specific noise profiles and annotation-error maps for both categorical and structured labels (e.g., segmentation masks).
Existing label-noise detection methods suffer from task specificity, limited learnability, and poor generalization across vision tasks. To address this, we propose the first unified, learnable framework for label-noise detection applicable to object detection, semantic segmentation, and instance segmentation. Our core innovation is an “error-correcting-via-error” paradigm: we reformulate label-noise detection as an instance segmentation problem by injecting controllable synthetic errors, and introduce a composite input representation alongside multi-task joint training. We establish a new authoritative benchmark on Cityscapes containing 459 real-world mislabeled instances. Extensive experiments demonstrate that our method significantly outperforms prior approaches, exhibiting strong generalization and robustness in both synthetic and real-world settings. This work provides a reproducible, scalable paradigm and empirical foundation for improving data quality in supervised learning.
This work addresses pervasive labeling errors—such as missing annotations, misclassifications, and imprecise bounding boxes—in object detection datasets. We propose REC✓D, a semi-automated correction framework that leverages pre-trained detectors to generate candidate mislabeling suggestions, employs lightweight crowdsourced micro-tasks for independent human verification by multiple annotators, and applies response aggregation to quantify annotation ambiguity and enhance correction robustness. To our knowledge, REC✓D is the first scalable, end-to-end system for detecting and correcting labels in object detection data. As a key contribution, we release a rigorously validated high-quality subset of pedestrian annotations from KITTI as a new benchmark. Experiments demonstrate that existing methods fail to detect up to 66% of ground-truth labeling errors, whereas REC✓D identifies and rectifies at least 24% of original annotation errors at a fraction of the cost of full manual re-annotation, substantially improving dataset quality and model evaluation reliability.
Widespread label noise (10–25%) in NLP benchmark datasets leads to systematic underestimation of model performance, with many purported “LLM failures” attributable to annotation errors rather than model limitations. Method: We propose LLM-as-a-judge—a framework leveraging ensemble judgments from GPT-4, Claude, and Llama, combined with consistency voting and error-sensitivity analysis to automatically detect mislabeled instances; we further apply label smoothing and confident learning for robust label recalibration. Contribution/Results: Comprehensive evaluation across the TRUE benchmark suite reveals substantial disparities in quality and efficiency among expert, crowdsourced, and LLM-generated annotations. After correction, state-of-the-art models achieve average accuracy gains of 3.2–7.8 percentage points. This work provides the first empirical evidence of systematic label-noise interference in LLM evaluation and introduces a scalable, collaborative adjudication paradigm that reframes data correction as model performance recalibration.
This study addresses the pervasive issues of semantic mislabeling and bounding box localization errors in object detection datasets by systematically evaluating, for the first time, the effectiveness of training-free feature-space methods for annotation error detection. Leveraging multiple pretrained embedding models, the approach is rigorously tested on both synthetic noise—including symmetric, asymmetric, and localization-type perturbations—and real-world annotation errors in the VOC2012 and KITTI datasets. Experimental results demonstrate that feature-space methods are highly effective at identifying semantic mislabels but exhibit limited capability in detecting localization inaccuracies. To facilitate future research, the authors publicly release all code and a curated set of verified erroneous annotations, establishing a valuable benchmark for the community.
To address the limitation of prior-dependent constraints undermining generalizability in hierarchical multi-label classification (HMC), this paper proposes EDR, a prior-free error-driven constraint discovery framework. EDR automatically identifies model misprediction patterns to induce interpretable, structured logical constraints; it further integrates a constraint-driven post-processing mechanism with a neuro-symbolic joint modeling architecture to jointly perform error detection, constraint recovery, and multi-level consistency verification. For the first time, EDR achieves fully automated, interpretable knowledge discovery and robust cross-domain constraint transfer without predefined constraints—enabling effective constraint learning even under label noise. Evaluated on multiple public benchmarks and a newly constructed military vehicle recognition dataset, EDR achieves an error detection F1-score exceeding 0.89 and constraint recovery accuracy above 92%, significantly improving both hierarchical consistency and overall classification performance.
This study investigates the impact of label noise on the generalization performance of language models, particularly in low-resource or high-noise settings. Focusing on Russian multi-domain text classification, it presents the first systematic comparison between Confident Learning and Dataset Cartography as automated methods for detecting label errors. Leveraging a fine-tuned rubert-base-cased model, the authors apply these techniques to filter noisy training data and validate their efficacy through controlled random-deletion baselines. Results demonstrate that Confident Learning substantially improves macro-F1 scores on small, high-noise datasets, whereas Dataset Cartography adopts a more conservative approach, removing fewer samples. Both methods consistently outperform random deletion, with their relative effectiveness closely dependent on dataset size and noise level.
Real-world data often contain label noise that severely degrades model performance. This work proposes Relabeler, a novel framework that, for the first time, jointly models both local and global relationships among data instances within a unified architecture to enable end-to-end label correction. By leveraging input features and observed labels in concert, Relabeler estimates the most probable clean labels through a data-centric learning paradigm that integrates relational modeling, probabilistic label inference, and end-to-end optimization. Extensive experiments demonstrate that the method consistently outperforms existing approaches across diverse datasets, noise types, and noise rates, achieving up to a 58% improvement in label correction accuracy and a 6% gain in downstream task performance.
Label noise in supervised classification can severely degrade model performance, necessitating label-cleaning methods that require no prior knowledge of the noise characteristics. This work proposes an unsupervised noise identification framework that constructs data subsets via Bernoulli random sampling and leverages the linear relationship between cross-validation error and subset noise level to build a mixture distribution for distinguishing clean and noisy samples. Theoretically, we prove that under this sampling scheme, the mean noise level converges to two separable distributions, offering the first probabilistic coupling analysis guarantee for unsupervised label cleaning. The method is classifier-agnostic and consistently achieves significant improvements in both noise detection accuracy and downstream classification performance across synthetic and real-world datasets.
This work addresses the gap between theory and practice in error detection and cleaning for tabular data by proposing and implementing an interactive web-based demonstration system that integrates error modeling, injection, and intelligent cleaning. For the first time, the system unifies machine learning–based data cleaning methods with dependency-aware error generation models within a visual platform, enabling users to upload tables, inject realistic errors, and interactively experience the cleaning process alongside interpretable explanations of the underlying mechanisms. By synergizing machine learning, database technologies, and web development, the project enhances intuitive understanding of error repair principles and provides a publicly accessible demo platform (https://cured.demo.calgo-lab.de/) that serves as an effective validation tool for both research and practical applications in data cleaning.
This study addresses the pervasive issue of data errors in real-world databases—such as missing values, redundancy, statistical biases, and outliers—which significantly degrade downstream analytical and machine learning performance. Recognizing that existing taxonomies are incomplete and terminology inconsistent, this work presents the first unified framework that integrates traditional data errors with statistically oriented inaccuracies critical in the AI era. It proposes a non-overlapping tripartite classification structure—comprising missing, erroneous, and redundant data—and systematically constructs a comprehensive catalog of 35 distinct error types. Through formal definitions, illustrative examples, and a thorough literature review, the paper establishes standardized terminology and precise characterizations, thereby offering a clear, rigorous theoretical foundation and practical toolkit for data quality assessment and cleaning.