Score
Designs and executes analyses and diagnostic pipelines that detect, quantify, and categorize incorrect, inconsistent, or unexpected labels in datasets; investigates discrepancies between assigned labels, annotator judgments, and available ground truth to identify root causes and patterns of labeling error. Builds audits, metrics, and correction or relabeling protocols to estimate the impact of label errors on model behavior and to guide remediation.
Real-world data often contain label noise that severely degrades model performance. This work proposes Relabeler, a novel framework that, for the first time, jointly models both local and global relationships among data instances within a unified architecture to enable end-to-end label correction. By leveraging input features and observed labels in concert, Relabeler estimates the most probable clean labels through a data-centric learning paradigm that integrates relational modeling, probabilistic label inference, and end-to-end optimization. Extensive experiments demonstrate that the method consistently outperforms existing approaches across diverse datasets, noise types, and noise rates, achieving up to a 58% improvement in label correction accuracy and a 6% gain in downstream task performance.
This paper addresses the reliable detection of post-deployment performance degradation (PDD) in unlabeled model-serving scenarios. We formally define the PDD monitoring task as distinguishing benign distributional shifts from genuine performance deterioration. To this end, we propose D3M—a label-free, gradient-free monitoring framework that leverages predictive disagreement across multiple models. We theoretically establish its low false-positive rate under non-degrading shifts and provide sample-complexity guarantees. By unifying theoretical analysis with empirical risk estimation, D3M achieves significant improvements over state-of-the-art baselines on standard benchmarks and a large-scale real-world internal medicine dataset. Our method delivers a verifiable, automated alerting mechanism for performance degradation in high-stakes machine learning systems.
Current CVML models struggle to localize and attribute failures at the subgroup level without labeled data. Method: We propose the first interactive error analysis framework integrating large language model (LLM) semantic understanding with visual analytics—requiring no manual annotations. It leverages CLIP-based semantic embeddings to cluster misclassified images into semantically coherent subgroups, then employs GPT-4 to generate interpretable, hypothesis-driven explanations. The framework further supports concept-guided interactive validation and comparative analysis. Contribution/Results: Evaluated on image classification, object detection, and semantic segmentation, our approach significantly improves both the efficiency and depth of error attribution—enabling an explainable shift from “where errors occur” to “why they occur.” Domain experts have validated its effectiveness and practical utility.
Label noise significantly degrades model performance, necessitating effective detection methods. This work proposes Adaptive Label Error Detection (ALED), which first extracts and denoises intermediate features from deep convolutional networks, then models each class as a multivariate Gaussian distribution on a low-dimensional manifold, and finally identifies mislabeled samples via a Bayesian likelihood ratio test. ALED is the first approach to integrate feature denoising, class-conditional Gaussian modeling, and likelihood ratio testing within a unified framework. Evaluated on multiple medical imaging datasets, it substantially outperforms existing techniques; fine-tuning models with labels corrected by ALED reduces test error rates by 33.8%, achieving both high sensitivity and precision.
Widespread label noise (10–25%) in NLP benchmark datasets leads to systematic underestimation of model performance, with many purported “LLM failures” attributable to annotation errors rather than model limitations. Method: We propose LLM-as-a-judge—a framework leveraging ensemble judgments from GPT-4, Claude, and Llama, combined with consistency voting and error-sensitivity analysis to automatically detect mislabeled instances; we further apply label smoothing and confident learning for robust label recalibration. Contribution/Results: Comprehensive evaluation across the TRUE benchmark suite reveals substantial disparities in quality and efficiency among expert, crowdsourced, and LLM-generated annotations. After correction, state-of-the-art models achieve average accuracy gains of 3.2–7.8 percentage points. This work provides the first empirical evidence of systematic label-noise interference in LLM evaluation and introduces a scalable, collaborative adjudication paradigm that reframes data correction as model performance recalibration.
Real-world datasets often suffer from multiple issues, including label noise, feature corruption, and spurious correlations, yet existing methods struggle to simultaneously identify erroneous samples and their specific error types with high precision. This work proposes DeMix, a novel framework that, for the first time, leverages influence vectors to characterize how individual training samples affect predictions on a validation set. By formulating data debugging as a multi-label classification task and incorporating intervention-based learning to extract invariant diagnostic criteria for each error type, DeMix enables joint identification of both corrupted samples and their underlying error categories. Evaluated across 11 benchmark tasks, DeMix improves the F1 score for data debugging by 22.61% on average and boosts downstream model performance by 9.32% after data repair, substantially outperforming current state-of-the-art approaches.
This work addresses a critical limitation in existing reconstruction-based methods for learning with noisy labels: their tendency to jointly assess the reliability of observed labels and pseudo-targets, which often leads to unreliable signals substituting one another and hinders effective denoising. To overcome this, the paper proposes TRACE, a novel framework that decouples the reliability evaluation of these two sources for the first time. Specifically, it evaluates observed labels through loss fitting, shallow-feature relational stability, and prediction consistency, while assessing pseudo-targets via model confidence. These independent reliability estimates are then used to separately govern label correction and sample reweighting. By preventing error propagation, TRACE generates more trustworthy pseudo-supervision, significantly outperforming current reconstruction-based approaches across multiple synthetic and real-world noisy benchmarks, and thereby enhancing model robustness and generalization.
为解决软件问题报告手动标注耗时费力的问题,提出LabelMate框架,利用大型语言模型自动生成并自动标注项目特定标签,无需预标注数据,提高标注准确率。
This work addresses the gap between theory and practice in error detection and cleaning for tabular data by proposing and implementing an interactive web-based demonstration system that integrates error modeling, injection, and intelligent cleaning. For the first time, the system unifies machine learning–based data cleaning methods with dependency-aware error generation models within a visual platform, enabling users to upload tables, inject realistic errors, and interactively experience the cleaning process alongside interpretable explanations of the underlying mechanisms. By synergizing machine learning, database technologies, and web development, the project enhances intuitive understanding of error repair principles and provides a publicly accessible demo platform (https://cured.demo.calgo-lab.de/) that serves as an effective validation tool for both research and practical applications in data cleaning.
This work addresses the challenge of reliably detecting label shift in small-batch scientific deployments, where target-domain labels are scarce. The authors propose a formal sequential testing framework that reframes model monitoring as an anytime-valid sequential hypothesis test. The key innovation lies in constructing a validation rule based on conditional e-values and non-negative martingales, which for the first time directly links the negative log predictive density difference to label shift detection. Theoretical guarantees are established by integrating likelihood ratios, Ville’s inequality, and Gaussian process regression. Simulations demonstrate that the method achieves strong finite-sample power and calibration robustness while rigorously controlling Type I error, significantly outperforming existing approaches that rely on re-estimating label distributions.