dataset audit and repair

Designs and implements processes, tools, and workflows to inspect and document datasets and benchmarks for annotation errors, missing or contradictory labels, underspecified or non-derivable items, visual or design inconsistencies, and compliance gaps. Builds and applies repair operations — including relabeling and curation procedures, dataset versioning, creation of controlled evaluation sets, and synthesize-and-verify pipelines — to correct identified problems and validate that the repaired dataset meets evaluation and compliance readiness criteria.

datasetauditandrepair

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.56
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

From Bugs to Benchmarks: A Comprehensive Survey of Software Defect Datasets

Apr 24, 2025
HZ
Hao-Nan Zhu
🏛️ University of California, Davis | University of Stuttgart

Large-scale, heterogeneous software defect datasets hinder efficient navigation and reuse by researchers. This paper systematically surveys 132 publicly available defect datasets and proposes a multidimensional evaluation framework—covering domain coverage, defect types, programming language distribution, construction methodologies, accessibility, and citation contexts—to achieve the first standardized metadata harmonization and empirical usability validation across a hundred-plus datasets. Through bibliometric analysis, citation network mapping, and cross-dimensional clustering, we identify test generation and automated program repair as the most widely supported application domains, while critical defect categories—including concurrency and security vulnerabilities—remain severely underrepresented. We further propose actionable dataset curation guidelines and a reusable assessment template, diagnosing pervasive issues such as incomplete coverage, disorganized structure, and outdated maintenance. Our work establishes a robust, empirically grounded data foundation for software defect detection, localization, repair, and AI-driven development.

Evaluating dataset scope, construction, availability, and usabilityIdentifying future research opportunities in defect dataset improvementSurveying 132 software defect datasets for comprehensive analysis

This work addresses the reliability of automatic program repair (APR) evaluation by proposing the first rigorous reproducibility criteria for APR-oriented defect datasets. Applying these standards—encompassing automated test execution, static analysis, patch behavior comparison, and test suite adequacy checks—to the widely used Defects4J benchmark reveals significant shortcomings: 21.6% of its defects are unsuitable for APR evaluation, and an additional 7.1% suffer from substantially inadequate test suites, collectively rendering 28.7% (239 defects) prone to unreliable assessment. To promote more rigorous future research, the authors release the first open-source Java APR evaluation framework and call on the community to prioritize benchmark quality in empirical studies.

automated program repairbenchmark datasetDefects4J

Automatic Data Repair: Are We Ready to Deploy?

Oct 01, 2023
WN
Wei Ni
🏛️ Zhejiang University | City University of Hong Kong

In the era of generative AI, dirty data severely degrades downstream analytical accuracy and model performance, yet existing automated data repair algorithms lack systematic benchmarking and practical deployment guidance. Method: We conduct a comprehensive benchmark study of 12 state-of-the-art repair algorithms across 12 diverse datasets under varying error rates and error types; propose an information-theoretic algorithm taxonomy; design novel evaluation metrics balancing practical utility and interpretability; and assess repair efficacy across four downstream tasks—statistical analysis, model training, anomaly detection, and generative AI input quality. Contribution/Results: We reveal the critical insight that “clean data does not imply optimal analytical performance.” Repair consistently improves downstream task outcomes; our heuristic optimization strategy boosts the average error reduction rate of SOTA methods by 23.6%; and we deliver an industrial-grade applicability guide alongside a curated list of open challenges.

Developing a unified optimization strategy to improve existing repair methodsEvaluating 12 data repair algorithms' performance under varying error conditionsProviding practical guidelines for deploying repair algorithms in real-world scenarios

Latest Papers

What's happening recently
View more

This work addresses the gap between theory and practice in error detection and cleaning for tabular data by proposing and implementing an interactive web-based demonstration system that integrates error modeling, injection, and intelligent cleaning. For the first time, the system unifies machine learning–based data cleaning methods with dependency-aware error generation models within a visual platform, enabling users to upload tables, inject realistic errors, and interactively experience the cleaning process alongside interpretable explanations of the underlying mechanisms. By synergizing machine learning, database technologies, and web development, the project enhances intuitive understanding of error repair principles and provides a publicly accessible demo platform (https://cured.demo.calgo-lab.de/) that serves as an effective validation tool for both research and practical applications in data cleaning.

data cleaningerror detectionerror models

This study addresses the frequent violations of clinical coding standards—such as ICD-10, CPT, and HL7 FHIR—by large language models when generating structured medical data, which impedes integration with electronic health record systems. To mitigate this, the authors propose and validate a closed-loop verification-and-repair framework that automatically detects and iteratively corrects formatting errors. The approach is evaluated using three open-source models—Qwen2.5-7B, Llama3.1-8B, and Gemma2-9B—deployed locally across 320 clinical scenarios. Results demonstrate a substantial improvement in schema compliance across all models, achieving an overall adherence rate of 99.0% and increasing individual model performance by 7.8 to 12.5 percentage points. Notably, 96% of detected errors were attributable to repairable representation-layer issues, with most resolved within one or two correction rounds, effectively compensating for the models’ limited understanding of healthcare IT standards.

clinical LLMshealthcare interoperabilityschema compliance

This work addresses the challenge of repairing syntactically erroneous markup in scientific and technical documents, a task for which existing evaluation lacks benchmarks grounded in real-world error patterns. To bridge this gap, we introduce TeXFix-Bench, the first multi-format (LaTeX/Typst/Markdown) repair benchmark built upon an empirically derived fault taxonomy. This taxonomy is constructed via grounded theory analysis of community-reported failures, and the benchmark incorporates AST-aware DocMut mutation operators to generate more challenging repair instances. Evaluating seven large language models under a unified zero-shot protocol across 48,651 attempts, we find that DocMut-induced errors are significantly harder to repair than those from conventional mutators, that Typst presents notably higher difficulty, and that 13.6–18.5% of syntactically correct repairs substantially alter the original document content—revealing that compilation success alone substantially overestimates true repair quality.

document repairfault taxonomyLaTeX

Hot Scholars

WJ

Wenbo Jiang

University of Electronic Science and Technology of China
AI securityBackdoor attack
KX

Kaidi Xu

Associate Professor, City University of Hong Kong
AI SecurityUncertainty QuantificationFormal Verification
PZ

Pan Zhou

Professor, School of Cyber Science and Engineering, Huazhong University and Science and Technology
Multimodal AI&LLMs,AI Security
SY

Sibei Yang

Associate Professor, School of Computer Science and Engineering, Sun Yat-Sen University
IS

Ilia Shumailov

AI Sequrity Company
Machine LearningComputer SecurityAdversarial Machine LearningAI Security