Score
Design and execute systematic experiments to benchmark and compare data cleaning algorithms and methods by building evaluation datasets or noise models, implementing metrics and reproducible protocols to measure correctness, coverage, runtime, and robustness. Analyze experimental results to identify common failure modes, quantify trade‑offs between approaches, and derive practical guidelines for selecting or improving cleaning methods.
A significant gap exists between theoretical definitions of data quality dimensions—such as accuracy, completeness, consistency, and timeliness—as stipulated in standards (e.g., ISO/IEC 25012) and the actual functionalities implemented in widely used data quality tools, with no systematic mapping analysis to date. Method: We conducted a systematic literature review and performed functional reverse engineering on seven mainstream open-source data quality tools. Contribution/Results: We present the first many-to-many mapping framework linking data quality dimensions to concrete tool capabilities, introducing a cross-dimensional, fine-grained functional categorization schema and generating a structured correspondence matrix covering all core dimensions. This work bridges the theory–practice divide, delivers an actionable guide for data quality assessment, and substantially enhances the scientific rigor and reusability of tool selection, functionality design, and standard implementation.
This study addresses a critical gap in data cleaning research—the lack of large-scale, real-world dirty datasets that hinder the effective evaluation of methods in practical settings. To bridge this gap, the authors construct the first large-scale dirty dataset comprising real postal addresses paired with their ground-truth counterparts. Leveraging this dataset, they conduct a systematic benchmarking study of state-of-the-art data cleaning approaches. Their experiments reveal substantial limitations of current methods when applied to real-world data, underscoring the need for more robust and context-aware techniques. The dataset and empirical findings not only establish a reliable benchmark for future research but also provide actionable insights to guide the development of data cleaning solutions tailored to real-world applications.
To address the challenge of balancing efficiency and model performance in data cleaning under resource constraints, this paper proposes a progressive cleaning optimization framework designed to maximize machine learning effectiveness. The method integrates error sensitivity analysis, incremental model evaluation, and a greedy selection strategy to dynamically recommend—per iteration—the most beneficial features to clean first. It supports adaptive handling of multiple ML algorithms and diverse error types, overcoming limitations of static cleaning pipelines and heuristic approaches based solely on feature importance. Experiments across multiple real-world datasets and mainstream ML models demonstrate an average prediction accuracy improvement of 5 percentage points, with gains up to 52 percentage points, significantly outperforming existing baselines. The core contribution is the first formulation of cleaning decisions as a sequence optimization problem explicitly targeting end-to-end model performance gain, enabling scalable, interpretable, and real-time cleaning recommendations.
Data cleaning remains highly manual, inefficient, and error-prone. This paper proposes the first goal-driven LLM-based framework for automatic workflow generation: given a dirty table and a target query, it end-to-end generates a minimal viable clean table along with executable cleaning steps—including deduplication, missing-value imputation, and format standardization. Our contributions are threefold: (1) We introduce the first benchmark dataset comprising annotated quadruples of (goal, dirty table, cleaning workflow, cleaned answer); (2) We design a zero-shot, multi-stage prompting framework—requiring no fine-tuning—that decomposes the task into goal column identification, data quality diagnosis, and operation-parameter generation; (3) We empirically validate that off-the-shelf LLMs possess inherent reasoning capabilities sufficient to generate high-quality, executable cleaning workflows across three major LLM families, significantly reducing human intervention.
To address the prevalence of noisy and low-quality training data in automatic code comment update tasks, this paper proposes a semantic- and overlap-aware data cleaning method. It jointly models code-comment semantic similarity—computed via BERT embeddings—and lexical overlap to produce a holistic cleaning score; an unsupervised tail-filtering strategy then automatically removes low-quality samples based on the score distribution. This is the first work to synergistically integrate semantic consistency and textual overlap for comment update data cleaning, eliminating reliance on manual annotations. After filtering over 30% of samples across three mainstream datasets, downstream models consistently achieve improved performance across all evaluation metrics. Human evaluation confirms significantly higher noise identification accuracy compared to random sampling. Moreover, models trained on cleaned data demonstrate enhanced generalization, as evidenced by improved performance on cleaned test sets.
This work addresses the gap between theory and practice in error detection and cleaning for tabular data by proposing and implementing an interactive web-based demonstration system that integrates error modeling, injection, and intelligent cleaning. For the first time, the system unifies machine learning–based data cleaning methods with dependency-aware error generation models within a visual platform, enabling users to upload tables, inject realistic errors, and interactively experience the cleaning process alongside interpretable explanations of the underlying mechanisms. By synergizing machine learning, database technologies, and web development, the project enhances intuitive understanding of error repair principles and provides a publicly accessible demo platform (https://cured.demo.calgo-lab.de/) that serves as an effective validation tool for both research and practical applications in data cleaning.
This work addresses common data quality issues in tabular data—such as missing values, inconsistent formatting, dependency violations, unit errors, and categorical ambiguities—by proposing a training-free cleaning approach. The method leverages few-shot annotations to guide large language models (LLMs) in reasoning about data flaws and automatically synthesizing reusable Python cleaning scripts equipped with guard conditions. Central to this approach is an evidence-based guarded repair mechanism that applies deterministic transformations only when specific dirty data patterns are detected and sufficiently supported by contextual evidence. By integrating data profiling, program synthesis, and cell-level validation, the framework ensures high logical precision. Evaluated on six benchmark datasets, the method achieves higher F1 scores than state-of-the-art rule-based, learning-based, and LLM-driven baselines on five datasets, while substantially reducing computational overhead and LLM invocation costs during repeated cleaning runs.
Traditional data preparation methods face limitations in semantic understanding and generalization, struggling to meet the rapidly growing demand for application-ready data. This work systematically reviews the application of large language models (LLMs) in three core tasks—data cleaning, integration, and augmentation—and proposes a task-centered taxonomy that, for the first time, delineates the evolutionary trajectory of LLM-driven data preparation techniques. Through a comprehensive literature review, the study examines key technologies such as prompt engineering, agent-based architectures, and semantic matching, alongside prevailing datasets and evaluation metrics. It highlights LLMs’ strengths in enhancing generalization and semantic comprehension while identifying critical challenges related to computational cost, hallucination, scalability, and the lack of standardized evaluation frameworks. The paper concludes by outlining a roadmap for future research and development in this emerging field.
Multivariate time series (MTS) are frequently affected by co-occurring quality issues, such as missing values, outliers, and constraint violations, which significantly undermine downstream analytics. Existing cleaning approaches fix only a limited set of such issues, making them ill-suited for scenarios where multiple quality problems arise simultaneously. Furthermore, these methods commonly depend on the availability of ground truth data or domain-specific rules, both of which are rarely accessible in real-world applications. In this paper, we introduce \sys, an agent system with reinforcement learning designed to clean multiple data quality issues in MTS. We cast the cleaning process as a joint optimization problem that simultaneously handles quality issue order and cleaning model selection, allowing efficient navigation of the large space of possible cleaning pipelines. Our framework relies on a hierarchical agent architecture, where a high-level agent determines the order in which data quality issues should be processed, while a low-level agent identifies the most suitable cleaning method for each issue. To guide the agent toward an optimal cleaning pipeline, we propose a dual-stage reward mechanism that couples upstream (cleaning) and downstream performance, enabling effective optimization without relying on ground truth. Our experimental results show that \sys consistently outperforms existing methods, achieving up to 96\% improvement in data cleaning quality and 27\% improvement in downstream performance.
Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.