Score
Designs and implements processes, tools, and workflows to inspect and document datasets and benchmarks for annotation errors, missing or contradictory labels, underspecified or non-derivable items, visual or design inconsistencies, and compliance gaps. Builds and applies repair operations — including relabeling and curation procedures, dataset versioning, creation of controlled evaluation sets, and synthesize-and-verify pipelines — to correct identified problems and validate that the repaired dataset meets evaluation and compliance readiness criteria.
Existing dataset documentation tools struggle to achieve real-world adoption due to ambiguous value propositions, misalignment with practical contexts, insufficient attention to human labor costs, and a lack of systemic integration. This study addresses these challenges through a mixed-methods systematic scoping review of 59 relevant publications, combining qualitative coding with quantitative synthesis to uncover the underlying motivations driving tool design and their relationship to institutional norms. The analysis identifies four key patterns that hinder adoption and advances a responsible AI design perspective that shifts emphasis from individual accountability to institutional solutions. The work advocates embedding sustainable documentation practices within organizational workflows and cultures, offering the HCI community actionable pathways toward institutionalizing responsible data stewardship.
Large-scale, heterogeneous software defect datasets hinder efficient navigation and reuse by researchers. This paper systematically surveys 132 publicly available defect datasets and proposes a multidimensional evaluation framework—covering domain coverage, defect types, programming language distribution, construction methodologies, accessibility, and citation contexts—to achieve the first standardized metadata harmonization and empirical usability validation across a hundred-plus datasets. Through bibliometric analysis, citation network mapping, and cross-dimensional clustering, we identify test generation and automated program repair as the most widely supported application domains, while critical defect categories—including concurrency and security vulnerabilities—remain severely underrepresented. We further propose actionable dataset curation guidelines and a reusable assessment template, diagnosing pervasive issues such as incomplete coverage, disorganized structure, and outdated maintenance. Our work establishes a robust, empirically grounded data foundation for software defect detection, localization, repair, and AI-driven development.
This work addresses the reliability of automatic program repair (APR) evaluation by proposing the first rigorous reproducibility criteria for APR-oriented defect datasets. Applying these standards—encompassing automated test execution, static analysis, patch behavior comparison, and test suite adequacy checks—to the widely used Defects4J benchmark reveals significant shortcomings: 21.6% of its defects are unsuitable for APR evaluation, and an additional 7.1% suffer from substantially inadequate test suites, collectively rendering 28.7% (239 defects) prone to unreliable assessment. To promote more rigorous future research, the authors release the first open-source Java APR evaluation framework and call on the community to prioritize benchmark quality in empirical studies.
In the era of generative AI, dirty data severely degrades downstream analytical accuracy and model performance, yet existing automated data repair algorithms lack systematic benchmarking and practical deployment guidance. Method: We conduct a comprehensive benchmark study of 12 state-of-the-art repair algorithms across 12 diverse datasets under varying error rates and error types; propose an information-theoretic algorithm taxonomy; design novel evaluation metrics balancing practical utility and interpretability; and assess repair efficacy across four downstream tasks—statistical analysis, model training, anomaly detection, and generative AI input quality. Contribution/Results: We reveal the critical insight that “clean data does not imply optimal analytical performance.” Repair consistently improves downstream task outcomes; our heuristic optimization strategy boosts the average error reduction rate of SOTA methods by 23.6%; and we deliver an industrial-grade applicability guide alongside a curated list of open challenges.
This work addresses the gap between theory and practice in error detection and cleaning for tabular data by proposing and implementing an interactive web-based demonstration system that integrates error modeling, injection, and intelligent cleaning. For the first time, the system unifies machine learning–based data cleaning methods with dependency-aware error generation models within a visual platform, enabling users to upload tables, inject realistic errors, and interactively experience the cleaning process alongside interpretable explanations of the underlying mechanisms. By synergizing machine learning, database technologies, and web development, the project enhances intuitive understanding of error repair principles and provides a publicly accessible demo platform (https://cured.demo.calgo-lab.de/) that serves as an effective validation tool for both research and practical applications in data cleaning.
This study addresses the frequent violations of clinical coding standards—such as ICD-10, CPT, and HL7 FHIR—by large language models when generating structured medical data, which impedes integration with electronic health record systems. To mitigate this, the authors propose and validate a closed-loop verification-and-repair framework that automatically detects and iteratively corrects formatting errors. The approach is evaluated using three open-source models—Qwen2.5-7B, Llama3.1-8B, and Gemma2-9B—deployed locally across 320 clinical scenarios. Results demonstrate a substantial improvement in schema compliance across all models, achieving an overall adherence rate of 99.0% and increasing individual model performance by 7.8 to 12.5 percentage points. Notably, 96% of detected errors were attributable to repairable representation-layer issues, with most resolved within one or two correction rounds, effectively compensating for the models’ limited understanding of healthcare IT standards.
This work addresses the challenge of repairing syntactically erroneous markup in scientific and technical documents, a task for which existing evaluation lacks benchmarks grounded in real-world error patterns. To bridge this gap, we introduce TeXFix-Bench, the first multi-format (LaTeX/Typst/Markdown) repair benchmark built upon an empirically derived fault taxonomy. This taxonomy is constructed via grounded theory analysis of community-reported failures, and the benchmark incorporates AST-aware DocMut mutation operators to generate more challenging repair instances. Evaluating seven large language models under a unified zero-shot protocol across 48,651 attempts, we find that DocMut-induced errors are significantly harder to repair than those from conventional mutators, that Typst presents notably higher difficulty, and that 13.6–18.5% of syntactically correct repairs substantially alter the original document content—revealing that compilation success alone substantially overestimates true repair quality.