classify error types

Designs and implements models, heuristics, or pipelines that assign one or more error-type labels to individual examples, producing per-sample scores or probabilities for categories such as label errors, feature errors, and spurious-error sources. Builds analysis and diagnostic tooling that supports multi-label error diagnosis and prioritization, routing examples to targeted downstream repair or evaluation actions.

classifyerrortypes

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Real-world datasets often suffer from multiple issues, including label noise, feature corruption, and spurious correlations, yet existing methods struggle to simultaneously identify erroneous samples and their specific error types with high precision. This work proposes DeMix, a novel framework that, for the first time, leverages influence vectors to characterize how individual training samples affect predictions on a validation set. By formulating data debugging as a multi-label classification task and incorporating intervention-based learning to extract invariant diagnostic criteria for each error type, DeMix enables joint identification of both corrupted samples and their underlying error categories. Evaluated across 11 benchmark tasks, DeMix improves the F1 score for data debugging by 22.61% on average and boosts downstream model performance by 9.32% after data repair, substantially outperforming current state-of-the-art approaches.

data qualityerror type identificationinfluence vectors

To address the limitation of prior-dependent constraints undermining generalizability in hierarchical multi-label classification (HMC), this paper proposes EDR, a prior-free error-driven constraint discovery framework. EDR automatically identifies model misprediction patterns to induce interpretable, structured logical constraints; it further integrates a constraint-driven post-processing mechanism with a neuro-symbolic joint modeling architecture to jointly perform error detection, constraint recovery, and multi-level consistency verification. For the first time, EDR achieves fully automated, interpretable knowledge discovery and robust cross-domain constraint transfer without predefined constraints—enabling effective constraint learning even under label noise. Evaluated on multiple public benchmarks and a newly constructed military vehicle recognition dataset, EDR achieves an error detection F1-score exceeding 0.89 and constraint recovery accuracy above 92%, significantly improving both hierarchical consistency and overall classification performance.

Detects errors in hierarchical multi-label classification without prior constraintsLearns explainable rules for machine learning model failure modesRecovers constraints for neurosymbolic models from error detection rules

Current CVML models struggle to localize and attribute failures at the subgroup level without labeled data. Method: We propose the first interactive error analysis framework integrating large language model (LLM) semantic understanding with visual analytics—requiring no manual annotations. It leverages CLIP-based semantic embeddings to cluster misclassified images into semantically coherent subgroups, then employs GPT-4 to generate interpretable, hypothesis-driven explanations. The framework further supports concept-guided interactive validation and comparative analysis. Contribution/Results: Evaluated on image classification, object detection, and semantic segmentation, our approach significantly improves both the efficiency and depth of error attribution—enabling an explainable shift from “where errors occur” to “why they occur.” Domain experts have validated its effectiveness and practical utility.

Enables interactive error analysis through subgroup discovery and validationIdentifies CVML model errors at subgroup level without labelsLeverages foundation models for semantic error interpretation

In industrial quality inspection, anomaly detection suffers from poor robustness due to high noise levels and sparse defective samples. To address this, we propose Iterative Refinement of Pseudo-labels (IRP), a self-supervised method that alternately evaluates sample credibility and removes misleading instances under feature-space consistency constraints—effectively purifying the training set dynamically without human annotations and generating high-fidelity self-supervised signals. IRP introduces the novel paradigm of “iterative data refinement,” significantly enhancing model robustness against label noise and cross-domain generalization capability. Evaluated on KSDD2 and MVTec AD benchmarks, IRP consistently outperforms existing unsupervised and self-supervised methods. Notably, under high-noise conditions, it achieves substantial improvements in detection accuracy and reduces false positive rates by over 25%.

Enhances defect detection accuracy in industrial quality control.Improves model performance by removing misleading data points.Outperforms traditional models in noisy industrial environments.

Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performance

Oct 24, 2024
ON
Omer Nahum
🏛️ Technion - Institute of Technology | Google Research

Widespread label noise (10–25%) in NLP benchmark datasets leads to systematic underestimation of model performance, with many purported “LLM failures” attributable to annotation errors rather than model limitations. Method: We propose LLM-as-a-judge—a framework leveraging ensemble judgments from GPT-4, Claude, and Llama, combined with consistency voting and error-sensitivity analysis to automatically detect mislabeled instances; we further apply label smoothing and confident learning for robust label recalibration. Contribution/Results: Comprehensive evaluation across the TRUE benchmark suite reveals substantial disparities in quality and efficiency among expert, crowdsourced, and LLM-generated annotations. After correction, state-of-the-art models achieve average accuracy gains of 3.2–7.8 percentage points. This work provides the first empirical evidence of systematic label-noise interference in LLM evaluation and introduces a scalable, collaborative adjudication paradigm that reframes data correction as model performance recalibration.

Comparing annotation quality from experts, crowdsourcing, and LLMsDetecting label errors in NLP benchmark datasetsMitigating mislabeled data effects on model performance

Latest Papers

What's happening recently
View more

This study addresses the limitations of existing code error detection methods—namely, the absence of large-scale datasets, insufficient capability for multi-error analysis, and lack of a unified classification framework—which hinder context-sensitive debugging in educational settings. To bridge this gap, the authors propose a three-tier hierarchical error taxonomy grounded in Python’s official exception hierarchy and introduce PyMETA, the first large-scale, fine-grained annotated dataset of student code errors, comprising 48,646 submissions and 97 expert-annotated multi-error samples. Leveraging static analysis, expert annotation, and prompt engineering, the work systematically evaluates fine-tuned models (e.g., CodeBERT) against prominent large language models (e.g., GPT-3.5, Gemini 2.5 Pro) across multiple error recognition tasks. Experiments show that Gemini 2.5 Pro achieves a macro F1 score of 81.8% under a multi-error “inclusion” criterion, yet prompt-based LLMs generally underperform fine-tuned smaller models and exhibit a consistent tendency to over-predict logical errors.

code error classificationeducational debugginghierarchical taxonomy

This study addresses the lag in benchmark development for multimodal large language models and the misalignment between diagnostic objectives and generated content in automated evaluation. To this end, we propose a multi-agent collaborative framework governed by a central Harness mechanism. This approach introduces a novel Harness governance paradigm that integrates sparse error taxonomy specification mapping to construct chart question-answering diagnostic samples on demand, ensuring strict alignment with externally specified diagnostic objectives during data generation. Experimental results demonstrate that 86.4% of the generated samples satisfy the target requirements, effectively revealing capability disparities across different models. These findings validate both the feasibility and efficacy of targeted data generation for model evaluation.

Benchmark ConstructionCapability Gap DiagnosisChart Question Answering

Hot Scholars

EC

Erik Cambria

Professor @ NTU CCDS & Visiting @ MIT Media Lab
Neurosymbolic AIMultimodal InteractionNLPAffective Computing
UB

Ulas Bagci

Northwestern University
artificial intelligencedeep learningbiomedical image analysismedical image computing
RS

Rohit Saxena

University of Edinburgh
Natural Language ProcessingMultimodal Machine Learning
AP

Aryo Pradipta Gema

University of Edinburgh
Language ModellingLLM HallucinationsClinical NLPKnowledge Graph
PM

Pasquale Minervini

University of Edinburgh, Miniml.AI, ELLIS Scholar
Generative AIMachine LearningNatural Language ProcessingMachine Reasoning