evaluate data cleaning methods

Design and execute systematic experiments to benchmark and compare data cleaning algorithms and methods by building evaluation datasets or noise models, implementing metrics and reproducible protocols to measure correctness, coverage, runtime, and robustness. Analyze experimental results to identify common failure modes, quantify trade‑offs between approaches, and derive practical guidelines for selecting or improving cleaning methods.

evaluatedatacleaningmethods

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.47
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses a critical gap in data cleaning research—the lack of large-scale, real-world dirty datasets that hinder the effective evaluation of methods in practical settings. To bridge this gap, the authors construct the first large-scale dirty dataset comprising real postal addresses paired with their ground-truth counterparts. Leveraging this dataset, they conduct a systematic benchmarking study of state-of-the-art data cleaning approaches. Their experiments reveal substantial limitations of current methods when applied to real-world data, underscoring the need for more robust and context-aware techniques. The dataset and empirical findings not only establish a reliable benchmark for future research but also provide actionable insights to guide the development of data cleaning solutions tailored to real-world applications.

benchmarkingdata cleaningdirty data

Step-by-Step Data Cleaning Recommendations to Improve ML Prediction Accuracy

Mar 14, 2025
SM
Sedir Mohammed
🏛️ Hasso Plattner Institute | University of Potsdam | University of Amsterdam

To address the challenge of balancing efficiency and model performance in data cleaning under resource constraints, this paper proposes a progressive cleaning optimization framework designed to maximize machine learning effectiveness. The method integrates error sensitivity analysis, incremental model evaluation, and a greedy selection strategy to dynamically recommend—per iteration—the most beneficial features to clean first. It supports adaptive handling of multiple ML algorithms and diverse error types, overcoming limitations of static cleaning pipelines and heuristic approaches based solely on feature importance. Experiments across multiple real-world datasets and mainstream ML models demonstrate an average prediction accuracy improvement of 5 percentage points, with gains up to 52 percentage points, significantly outperforming existing baselines. The core contribution is the first formulation of cleaning decisions as a sequence optimization problem explicitly targeting end-to-end model performance gain, enabling scalable, interpretable, and real-time cleaning recommendations.

Optimizes data cleaning to improve ML prediction accuracy.Provides step-by-step recommendations for feature cleaning priorities.Reduces costs by efficient resource allocation in cleaning.

AutoDCWorkflow: LLM-based Data Cleaning Workflow Auto-Generation and Benchmark

Dec 09, 2024
LL
Lan Li
🏛️ University of Illinois, Urbana-Champaign

Data cleaning remains highly manual, inefficient, and error-prone. This paper proposes the first goal-driven LLM-based framework for automatic workflow generation: given a dirty table and a target query, it end-to-end generates a minimal viable clean table along with executable cleaning steps—including deduplication, missing-value imputation, and format standardization. Our contributions are threefold: (1) We introduce the first benchmark dataset comprising annotated quadruples of (goal, dirty table, cleaning workflow, cleaned answer); (2) We design a zero-shot, multi-stage prompting framework—requiring no fine-tuning—that decomposes the task into goal column identification, data quality diagnosis, and operation-parameter generation; (3) We empirically validate that off-the-shelf LLMs possess inherent reasoning capabilities sufficient to generate high-quality, executable cleaning workflows across three major LLM families, significantly reducing human intervention.

Addressing format inconsistencies, type errors, duplicates in datasetsAutomating data cleaning workflow generation using LLMsEvaluating workflow quality against human-curated benchmarks

CupCleaner: A Data Cleaning Approach for Comment Updating

Aug 14, 2023
QL
Qingyuan Liang
🏛️ Peking University | Institute of Software, Chinese Academy of Sciences

To address the prevalence of noisy and low-quality training data in automatic code comment update tasks, this paper proposes a semantic- and overlap-aware data cleaning method. It jointly models code-comment semantic similarity—computed via BERT embeddings—and lexical overlap to produce a holistic cleaning score; an unsupervised tail-filtering strategy then automatically removes low-quality samples based on the score distribution. This is the first work to synergistically integrate semantic consistency and textual overlap for comment update data cleaning, eliminating reliance on manual annotations. After filtering over 30% of samples across three mainstream datasets, downstream models consistently achieve improved performance across all evaluation metrics. Human evaluation confirms significantly higher noise identification accuracy compared to random sampling. Moreover, models trained on cleaned data demonstrate enhanced generalization, as evidenced by improved performance on cleaned test sets.

Cleans noisy comment update datasets for software evolutionEnhances code-comment consistency through hybrid statistical approachImproves model training by filtering inconsistent data samples

Latest Papers

What's happening recently
View more

This work addresses the gap between theory and practice in error detection and cleaning for tabular data by proposing and implementing an interactive web-based demonstration system that integrates error modeling, injection, and intelligent cleaning. For the first time, the system unifies machine learning–based data cleaning methods with dependency-aware error generation models within a visual platform, enabling users to upload tables, inject realistic errors, and interactively experience the cleaning process alongside interpretable explanations of the underlying mechanisms. By synergizing machine learning, database technologies, and web development, the project enhances intuitive understanding of error repair principles and provides a publicly accessible demo platform (https://cured.demo.calgo-lab.de/) that serves as an effective validation tool for both research and practical applications in data cleaning.

data cleaningerror detectionerror models

This work addresses common data quality issues in tabular data—such as missing values, inconsistent formatting, dependency violations, unit errors, and categorical ambiguities—by proposing a training-free cleaning approach. The method leverages few-shot annotations to guide large language models (LLMs) in reasoning about data flaws and automatically synthesizing reusable Python cleaning scripts equipped with guard conditions. Central to this approach is an evidence-based guarded repair mechanism that applies deterministic transformations only when specific dirty data patterns are detected and sufficiently supported by contextual evidence. By integrating data profiling, program synthesis, and cell-level validation, the framework ensures high logical precision. Evaluated on six benchmark datasets, the method achieves higher F1 scores than state-of-the-art rule-based, learning-based, and LLM-driven baselines on five datasets, while substantially reducing computational overhead and LLM invocation costs during repeated cleaning runs.

data qualityLLM-based cleaningreusable programs

Traditional data preparation methods face limitations in semantic understanding and generalization, struggling to meet the rapidly growing demand for application-ready data. This work systematically reviews the application of large language models (LLMs) in three core tasks—data cleaning, integration, and augmentation—and proposes a task-centered taxonomy that, for the first time, delineates the evolutionary trajectory of LLM-driven data preparation techniques. Through a comprehensive literature review, the study examines key technologies such as prompt engineering, agent-based architectures, and semantic matching, alongside prevailing datasets and evaluation metrics. It highlights LLMs’ strengths in enhancing generalization and semantic comprehension while identifying critical challenges related to computational cost, hallucination, scalability, and the lack of standardized evaluation frameworks. The paper concludes by outlining a roadmap for future research and development in this emerging field.

data cleaningdata enrichmentdata integration

Multivariate time series (MTS) are frequently affected by co-occurring quality issues, such as missing values, outliers, and constraint violations, which significantly undermine downstream analytics. Existing cleaning approaches fix only a limited set of such issues, making them ill-suited for scenarios where multiple quality problems arise simultaneously. Furthermore, these methods commonly depend on the availability of ground truth data or domain-specific rules, both of which are rarely accessible in real-world applications. In this paper, we introduce \sys, an agent system with reinforcement learning designed to clean multiple data quality issues in MTS. We cast the cleaning process as a joint optimization problem that simultaneously handles quality issue order and cleaning model selection, allowing efficient navigation of the large space of possible cleaning pipelines. Our framework relies on a hierarchical agent architecture, where a high-level agent determines the order in which data quality issues should be processed, while a low-level agent identifies the most suitable cleaning method for each issue. To guide the agent toward an optimal cleaning pipeline, we propose a dual-stage reward mechanism that couples upstream (cleaning) and downstream performance, enabling effective optimization without relying on ground truth. Our experimental results show that \sys consistently outperforms existing methods, achieving up to 96\% improvement in data cleaning quality and 27\% improvement in downstream performance.

data cleaningdata quality issuesmissing values

Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.

benchmarkingdataset qualitymodel datasets

Hot Scholars

XL

Xiaomeng Li

Assistant Professor, The Hong Kong University of Science and Technology
Medical Image AnalysisAI in HealthcareDeep Learning
YZ

Yuru Zhang

PhD Candidate of Computer Science, University of Nebraska-Lincoln
Wireless CommunicationMachine LearningEdge Computing
ST

Shu-Tao Xia

SIGS, Tsinghua University
coding and information theorymachine learningcomputer visionAI security
YF

Yuxuan Fan

Peking University
Natural Language Processing
MD

Muhammad Dehan Al Kautsar

Mohamed bin Zayed University of Artificial Intelligence
Natural Language ProcessingMultilingualityHuman-Centered NLP