iterative data refinement

Designs and implements iterative workflows and the observer-checker-corrector (OCC) framework that detect, check, and correct noisy, missing, or inconsistent annotations to produce high-fidelity grounded training datasets. This competence covers building automated validators and correctors, human-in-the-loop interfaces for high-density interactions, and iteration-aware metrics and processes to evaluate and converge data quality.

iterativedatarefinement

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$156K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Iterative Data Curation with Theoretical Guarantees

Oct 13, 2025
VY
Väinö Yrjänäinen
🏛️ Uppsala University | Chalmers University of Technology

Manual verification of large-scale dynamic datasets is infeasible, leading to challenges in ensuring data accuracy. Method: This paper proposes a theory-driven, iterative data cleaning framework integrating error detection, automated correction, and progressive optimization. It formally models the iterative process, conducts accuracy testing, and performs probabilistic convergence analysis—validated via simulations and real-world case studies. Contribution/Results: The framework establishes the first rigorous theoretical guarantee of *probabilistic convergence to zero errors* for iterative data cleaning and proves that error detection accelerates error decay. Empirical results demonstrate that it significantly outperforms baseline methods in accuracy improvement and progressively approaches a fully correct dataset state, thereby unifying theoretical soundness with practical efficacy.

Developing automated methods to improve accuracy in large datasetsEnsuring asymptotic elimination of all data errors with probability oneProviding theoretical guarantees for error reduction through iterative curation

This work addresses the critical bottleneck in automated program verification—synthesizing inductive loop invariants—where existing large language models often produce invalid or inefficient candidates on challenging instances. We introduce Wonda, a novel data curation pipeline that formally defines, for the first time, rigorous properties of high-quality invariants and constructs a refined training set through AST normalization and LLM-driven semantic rewriting. A small language model (4B parameters) fine-tuned on this curated data achieves, without any inference-time overhead, a twofold improvement in both correctness and speedup on hard InvBench instances, matching the performance of GPT-OSS-120B and approaching that of GPT-5.2, while boosting the virtual best solver’s verification performance by 14.2%.

automated reasoninginductive loop invariantsinvariant synthesis

From Label Error Detection to Correction: A Modular Framework and Benchmark for Object Detection Datasets

Aug 06, 2025
SP
Sarina Penquitt
🏛️ University of Wuppertal | Quality Match | Osnabrück University

This work addresses pervasive labeling errors—such as missing annotations, misclassifications, and imprecise bounding boxes—in object detection datasets. We propose REC✓D, a semi-automated correction framework that leverages pre-trained detectors to generate candidate mislabeling suggestions, employs lightweight crowdsourced micro-tasks for independent human verification by multiple annotators, and applies response aggregation to quantify annotation ambiguity and enhance correction robustness. To our knowledge, REC✓D is the first scalable, end-to-end system for detecting and correcting labels in object detection data. As a key contribution, we release a rigorously validated high-quality subset of pedestrian annotations from KITTI as a new benchmark. Experiments demonstrate that existing methods fail to detect up to 66% of ground-truth labeling errors, whereas REC✓D identifies and rectifies at least 24% of original annotation errors at a fraction of the cost of full manual re-annotation, substantially improving dataset quality and model evaluation reliability.

Address missing and inaccurate annotations in benchmark datasetsDetect and correct label errors in object detection datasetsImprove label quality through crowdsourced verification

Technical Report for Egocentric Mistake Detection for the HoloAssist Challenge

Jun 06, 2025
CP
Constantin Patsch
🏛️ Technical University of Munich

This work addresses the online detection of procedural and executional errors in first-person videos—specifically, procedural errors (e.g., step misordering) and executional errors (e.g., motion inaccuracies or tool misuse). We propose the first end-to-end online detection-feedback closed-loop framework that unifies modeling of both error types. Our method integrates temporal action recognition, sliding-window online inference, multimodal feature alignment, and leverages large language models to generate interpretable natural-language feedback. Unlike prior approaches targeting only one error category, ours enables fine-grained, real-time, and explainable joint detection and intervention for both error classes. Evaluated on the HoloAssist benchmark, our framework ranks second in the error detection task, demonstrating robustness and practical utility in real-world industrial and educational settings.

Detects procedural and execution errors in real-timeGenerates explanatory feedback using large language modelsValidates effectiveness on HoloAssist benchmark

Improving ML Training Data with Gold-Standard Quality Metrics

Dec 23, 2025
LB
Leslie Barrett
🏛️ Bloomberg LP | Google

Manual annotation suffers from inconsistent quality and lacks systematic evaluation. Method: This paper proposes a consensus-based quality measurement framework grounded in multi-round annotation statistics, using dynamic decay of inter-annotator agreement variance as the core metric—established here as a “gold standard” for data quality. Recognizing annotators’ significant warm-up period but prohibitive cost of full-sample multiple annotation, we design a low-redundancy, high-efficiency progressive annotation protocol. The approach integrates statistical consistency analysis, variance convergence modeling, and label confidence estimation. Contribution/Results: Our paradigm substantially enhances data quality’s measurability and controllability: experiments show 3.2–7.8% accuracy gains across multiple NLP tasks and over 30% reduction in annotation redundancy.

Collecting high-quality training data without requiring multiple tags per itemEnhancing data quality through iterative tagging to reduce varianceEvaluating hand-tagged training data quality using statistical consistency metrics

Latest Papers

What's happening recently
View more

This study addresses the risk of silent failures in AI scientific data reuse, where context loss can produce numerically valid yet semantically conflicting results. To overcome the limitations of traditional lineage and descriptive provenance tracking, this work proposes the concept of "domain boundaries," which explicitly encodes data generation conditions and binds them to standardized data containers, thereby enabling automated detection of invalid reuse scenarios. Experimental evaluation demonstrates that the proposed approach successfully identifies 17 out of 18 injected faults, significantly outperforming existing baseline methods. Furthermore, its detection performance remains robust regardless of dataset scale. These findings establish a new paradigm for ensuring safe cross-context data reuse in AI-driven scientific research.

AI-ready scientific datacontext lossdata reuse

This study addresses the pervasive issues of semantic mislabeling and bounding box localization errors in object detection datasets by systematically evaluating, for the first time, the effectiveness of training-free feature-space methods for annotation error detection. Leveraging multiple pretrained embedding models, the approach is rigorously tested on both synthetic noise—including symmetric, asymmetric, and localization-type perturbations—and real-world annotation errors in the VOC2012 and KITTI datasets. Experimental results demonstrate that feature-space methods are highly effective at identifying semantic mislabels but exhibit limited capability in detecting localization inaccuracies. To facilitate future research, the authors publicly release all code and a curated set of verified erroneous annotations, establishing a valuable benchmark for the community.

annotation errorsfeature-space methodsobject detection

Existing methods struggle to effectively evaluate the scientific validity of scientific images due to the weak correlation between perceptual quality metrics and scientific accuracy, as well as the limited domain-specific verification capabilities of general-purpose language models. This work proposes SIU²A, a novel framework that systematically defines scientific image utility—encompassing error detectability and correctability—and upgradability, which refers to restorability to scientific fidelity. The authors introduce SIU²A-Benchmark, a comprehensive dataset covering four categories of scientific distortions, along with a two-stage evaluation protocol that first assesses error identification and then evaluates restoration quality. Experimental results reveal significant shortcomings in current multimodal systems regarding both scientific error detection and faithful correction, highlighting a fundamental gap between visual perception and scientific usability.

AI-generated contenterror detectionimage integrity

Hot Scholars

RC

Renqi Chen

Southern University of Science and Technology, Fudan University
AI4ScienceLarge Language ModelMulti-Modal Language ModelMars Computing
YY

Yi Yuan

NetEase Fuxi AI Lab
deep learningcomputer vision
JL

Jianxun Lian

Microsoft Research Asia
LLM AgentAnthropomorphic IntelligenceUser ModelingRecommendation System
YL

Yiming Liao

Meta
Machine LearningRecommender SystemData Mining
XZ

Xinzhe Zheng

National University of Singapore
AI for BiomedicineAI for Science