Score
Designs and implements models, heuristics, or pipelines that assign one or more error-type labels to individual examples, producing per-sample scores or probabilities for categories such as label errors, feature errors, and spurious-error sources. Builds analysis and diagnostic tooling that supports multi-label error diagnosis and prioritization, routing examples to targeted downstream repair or evaluation actions.
A significant gap exists between academic research and industrial practice in debugging machine learning (ML) systems. Method: We propose the first comprehensive, lifecycle-spanning taxonomy of ML debugging faults and corresponding mitigation methods, derived from a systematic literature review (SLR), in-depth interviews with 28 ML practitioners, and empirical analysis of 1,247 GitHub issues. Contribution/Results: Our study identifies 13 core debugging challenges; only 48% are addressed by existing academic work, while 52.6% of GitHub issues and 70.3% of interview-elicited problems lack corresponding methodological support. Critically, we quantitatively demonstrate that over half of real-world ML debugging difficulties remain unaddressed by current research—revealing a substantial knowledge gap. This work establishes a foundational classification framework, provides empirical evidence of methodological coverage gaps, and delivers a prioritized roadmap to bridge the theory-practice divide in ML debugging.
Real-world datasets often suffer from multiple issues, including label noise, feature corruption, and spurious correlations, yet existing methods struggle to simultaneously identify erroneous samples and their specific error types with high precision. This work proposes DeMix, a novel framework that, for the first time, leverages influence vectors to characterize how individual training samples affect predictions on a validation set. By formulating data debugging as a multi-label classification task and incorporating intervention-based learning to extract invariant diagnostic criteria for each error type, DeMix enables joint identification of both corrupted samples and their underlying error categories. Evaluated across 11 benchmark tasks, DeMix improves the F1 score for data debugging by 22.61% on average and boosts downstream model performance by 9.32% after data repair, substantially outperforming current state-of-the-art approaches.
To address the limitation of prior-dependent constraints undermining generalizability in hierarchical multi-label classification (HMC), this paper proposes EDR, a prior-free error-driven constraint discovery framework. EDR automatically identifies model misprediction patterns to induce interpretable, structured logical constraints; it further integrates a constraint-driven post-processing mechanism with a neuro-symbolic joint modeling architecture to jointly perform error detection, constraint recovery, and multi-level consistency verification. For the first time, EDR achieves fully automated, interpretable knowledge discovery and robust cross-domain constraint transfer without predefined constraints—enabling effective constraint learning even under label noise. Evaluated on multiple public benchmarks and a newly constructed military vehicle recognition dataset, EDR achieves an error detection F1-score exceeding 0.89 and constraint recovery accuracy above 92%, significantly improving both hierarchical consistency and overall classification performance.
Current CVML models struggle to localize and attribute failures at the subgroup level without labeled data. Method: We propose the first interactive error analysis framework integrating large language model (LLM) semantic understanding with visual analytics—requiring no manual annotations. It leverages CLIP-based semantic embeddings to cluster misclassified images into semantically coherent subgroups, then employs GPT-4 to generate interpretable, hypothesis-driven explanations. The framework further supports concept-guided interactive validation and comparative analysis. Contribution/Results: Evaluated on image classification, object detection, and semantic segmentation, our approach significantly improves both the efficiency and depth of error attribution—enabling an explainable shift from “where errors occur” to “why they occur.” Domain experts have validated its effectiveness and practical utility.
In industrial quality inspection, anomaly detection suffers from poor robustness due to high noise levels and sparse defective samples. To address this, we propose Iterative Refinement of Pseudo-labels (IRP), a self-supervised method that alternately evaluates sample credibility and removes misleading instances under feature-space consistency constraints—effectively purifying the training set dynamically without human annotations and generating high-fidelity self-supervised signals. IRP introduces the novel paradigm of “iterative data refinement,” significantly enhancing model robustness against label noise and cross-domain generalization capability. Evaluated on KSDD2 and MVTec AD benchmarks, IRP consistently outperforms existing unsupervised and self-supervised methods. Notably, under high-noise conditions, it achieves substantial improvements in detection accuracy and reduces false positive rates by over 25%.
Widespread label noise (10–25%) in NLP benchmark datasets leads to systematic underestimation of model performance, with many purported “LLM failures” attributable to annotation errors rather than model limitations. Method: We propose LLM-as-a-judge—a framework leveraging ensemble judgments from GPT-4, Claude, and Llama, combined with consistency voting and error-sensitivity analysis to automatically detect mislabeled instances; we further apply label smoothing and confident learning for robust label recalibration. Contribution/Results: Comprehensive evaluation across the TRUE benchmark suite reveals substantial disparities in quality and efficiency among expert, crowdsourced, and LLM-generated annotations. After correction, state-of-the-art models achieve average accuracy gains of 3.2–7.8 percentage points. This work provides the first empirical evidence of systematic label-noise interference in LLM evaluation and introduces a scalable, collaborative adjudication paradigm that reframes data correction as model performance recalibration.
本文针对多标签文本分类中的校准问题,提出了一种新的分箱方案以准确估计校准误差,解决了现有方法低估误差或反映标签频率的问题。
研究利用治理记录作为监督,通过验证者选择的自训练方法修复结构化工作流,提高计划被接受的数量和效率。
This study addresses the limitations of existing code error detection methods—namely, the absence of large-scale datasets, insufficient capability for multi-error analysis, and lack of a unified classification framework—which hinder context-sensitive debugging in educational settings. To bridge this gap, the authors propose a three-tier hierarchical error taxonomy grounded in Python’s official exception hierarchy and introduce PyMETA, the first large-scale, fine-grained annotated dataset of student code errors, comprising 48,646 submissions and 97 expert-annotated multi-error samples. Leveraging static analysis, expert annotation, and prompt engineering, the work systematically evaluates fine-tuned models (e.g., CodeBERT) against prominent large language models (e.g., GPT-3.5, Gemini 2.5 Pro) across multiple error recognition tasks. Experiments show that Gemini 2.5 Pro achieves a macro F1 score of 81.8% under a multi-error “inclusion” criterion, yet prompt-based LLMs generally underperform fine-tuned smaller models and exhibit a consistent tendency to over-predict logical errors.
为解决软件问题报告手动标注耗时费力的问题,提出LabelMate框架,利用大型语言模型自动生成并自动标注项目特定标签,无需预标注数据,提高标注准确率。
This study addresses the lag in benchmark development for multimodal large language models and the misalignment between diagnostic objectives and generated content in automated evaluation. To this end, we propose a multi-agent collaborative framework governed by a central Harness mechanism. This approach introduces a novel Harness governance paradigm that integrates sparse error taxonomy specification mapping to construct chart question-answering diagnostic samples on demand, ensuring strict alignment with externally specified diagnostic objectives during data generation. Experimental results demonstrate that 86.4% of the generated samples satisfy the target requirements, effectively revealing capability disparities across different models. These findings validate both the feasibility and efficacy of targeted data generation for model evaluation.