data cleaning

Designs and implements processes, scripts, and pipelines to detect, correct, remove, or document errors, inconsistencies, duplicates, missing values, and formatting or type problems in datasets—including normalization, type conversion, deduplication, imputation, and outlier handling. Analyzes data quality, defines validation rules and cleaning heuristics, and produces reproducible, auditable cleaning workflows and tooling to prepare datasets for downstream analysis or modeling.

datacleaning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.84
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$216K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the pervasive issue of data errors in real-world databases—such as missing values, redundancy, statistical biases, and outliers—which significantly degrade downstream analytical and machine learning performance. Recognizing that existing taxonomies are incomplete and terminology inconsistent, this work presents the first unified framework that integrates traditional data errors with statistically oriented inaccuracies critical in the AI era. It proposes a non-overlapping tripartite classification structure—comprising missing, erroneous, and redundant data—and systematically constructs a comprehensive catalog of 35 distinct error types. Through formal definitions, illustrative examples, and a thorough literature review, the paper establishes standardized terminology and precise characterizations, thereby offering a clear, rigorous theoretical foundation and practical toolkit for data quality assessment and cleaning.

data errorsdata qualityerror taxonomy

DataLens: ML-Oriented Interactive Tabular Data Quality Dashboard

Jan 28, 2025
MA
Mohamed Abdelaal
🏛️ Software AG | TU Darmstadt

Existing data management tools suffer from limited automation, poor interactivity, and insufficient integration with ML workflows, compromising data quality and hindering analytical and modeling performance. To address this, we propose an interactive, ML-oriented tabular data quality dashboard that establishes an adaptive, human-in-the-loop + ML-driven data cleaning闭环. Our approach integrates data profiling, multi-strategy error detection and repair—including statistical analysis, rule-based engines, and supervised/semi-supervised models—while supporting expert rule validation and labeling. Cleaning strategies are iteratively refined using downstream model performance feedback. Furthermore, we unify DataSheets, MLflow, and Delta Lake to ensure reproducibility, traceability, and versioning of the cleaning pipeline. Experiments across multiple benchmark datasets demonstrate significant improvements: error identification rate and repair accuracy increase notably, downstream ML models achieve an average 7.2% accuracy gain, and cleaning time decreases by 40%.

Data ManagementData QualityMachine Learning Integration

This work addresses the gap between theory and practice in error detection and cleaning for tabular data by proposing and implementing an interactive web-based demonstration system that integrates error modeling, injection, and intelligent cleaning. For the first time, the system unifies machine learning–based data cleaning methods with dependency-aware error generation models within a visual platform, enabling users to upload tables, inject realistic errors, and interactively experience the cleaning process alongside interpretable explanations of the underlying mechanisms. By synergizing machine learning, database technologies, and web development, the project enhances intuitive understanding of error repair principles and provides a publicly accessible demo platform (https://cured.demo.calgo-lab.de/) that serves as an effective validation tool for both research and practical applications in data cleaning.

data cleaningerror detectionerror models

MechDetect: Detecting Data-Dependent Errors

Dec 03, 2025
PJ
Philipp Jung
🏛️ Berlin University of Applied Sciences and Technology

The core challenge in data quality monitoring lies in error provenance—specifically, identifying the underlying mechanisms that generate errors—a problem largely overlooked by existing work, which seldom models such mechanisms explicitly. This paper focuses on errors arising from intrinsic dependencies within data and proposes MechDetect, the first method to systematically extend missing-data mechanism detection to diverse error types—including outliers, inconsistencies, and format violations. Leveraging joint statistical modeling and supervised learning, MechDetect simultaneously models tabular data and their error masks to automatically determine whether observed errors stem from inherent characteristics of the original data. Extensive experiments across multiple benchmark datasets demonstrate that MechDetect significantly outperforms state-of-the-art baselines in accurately diagnosing error-generation mechanisms. By providing mechanistic interpretability, it establishes a theoretical foundation and practical framework for explainable data repair.

Detect data-dependent error generation mechanismsEstimate error dependency using machine learning modelsExtend missing value analysis to other error types

AutoDCWorkflow: LLM-based Data Cleaning Workflow Auto-Generation and Benchmark

Dec 09, 2024
LL
Lan Li
🏛️ University of Illinois, Urbana-Champaign

Data cleaning remains highly manual, inefficient, and error-prone. This paper proposes the first goal-driven LLM-based framework for automatic workflow generation: given a dirty table and a target query, it end-to-end generates a minimal viable clean table along with executable cleaning steps—including deduplication, missing-value imputation, and format standardization. Our contributions are threefold: (1) We introduce the first benchmark dataset comprising annotated quadruples of (goal, dirty table, cleaning workflow, cleaned answer); (2) We design a zero-shot, multi-stage prompting framework—requiring no fine-tuning—that decomposes the task into goal column identification, data quality diagnosis, and operation-parameter generation; (3) We empirically validate that off-the-shelf LLMs possess inherent reasoning capabilities sufficient to generate high-quality, executable cleaning workflows across three major LLM families, significantly reducing human intervention.

Addressing format inconsistencies, type errors, duplicates in datasetsAutomating data cleaning workflow generation using LLMsEvaluating workflow quality against human-curated benchmarks

Latest Papers

What's happening recently
View more

This study addresses a critical gap in data cleaning research—the lack of large-scale, real-world dirty datasets that hinder the effective evaluation of methods in practical settings. To bridge this gap, the authors construct the first large-scale dirty dataset comprising real postal addresses paired with their ground-truth counterparts. Leveraging this dataset, they conduct a systematic benchmarking study of state-of-the-art data cleaning approaches. Their experiments reveal substantial limitations of current methods when applied to real-world data, underscoring the need for more robust and context-aware techniques. The dataset and empirical findings not only establish a reliable benchmark for future research but also provide actionable insights to guide the development of data cleaning solutions tailored to real-world applications.

benchmarkingdata cleaningdirty data

This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.

data qualitydomain expertsno-code

This study addresses a critical limitation in traditional reproducible research, where sharing only code and results fails to expose the implicit assumptions, expectations, and premises underlying an analyst’s reasoning—thereby hindering thorough evaluation of analytical quality. To overcome this, the paper proposes a formal modeling framework that explicitly translates the analyst’s tacit reasoning process into structured logical representations, statically capturing the construction logic of the analysis. This approach enables systematic scrutiny of the analytical chain of reasoning, assumption sensitivity, and conclusion robustness—even in the absence of the original data. Empirical validation on representative data analysis tasks demonstrates the framework’s effectiveness, achieving both logical visualization and data-free static assessment of analytical integrity.

analysis reasoningassumptionsdata analysis

This work addresses the significant limitations of spreadsheet-based analysis in reproducibility, auditability, version control, and automation. It proposes a migration pathway from Excel to research-grade analytical workflows by leveraging Python’s pandas library as a bridge. The study introduces an innovative set of Excel-to-pandas mapping rules, categorizes nine canonical workflow patterns, and compiles a catalog of common failure modes. Seven end-to-end real-world examples demonstrate the approach in practice. By retaining Excel as a familiar interface for input and output while integrating version control, automated refreshing, and seamless incorporation of statistical and machine learning methods, the proposed framework enables governed, reproducible, and auditable tabular data analysis.

auditabilitydata analysisgovernance

Hot Scholars

DC

Dumitru-Clementin Cercel

Teaching Assistant of Computer Science, University Politehnica of Bucharest
Social Network AnalysisNatural Language ProcessingInformation RetrievalMachine Learning
LB

Lei Bai

Shanghai AI Laboratory
Foundation ModelScience IntelligenceMulti-Agent SystemAutonomous Discovery
JP

Jiaxin Pei

Stanford University, The University of Texas at Austin
Human-Centered AINLPHuman-Computer InteractionComputational Social Science
VG

Vivek Gupta

Assistant Professor of Computer Science, Arizona State University
Artificial IntelligenceNatural Language ProcessingLarge Language ModelsInformation Retrieval
RK

Rick Kazman

Professor, University of Hawaii
Software engineeringsoftware architecturetechnical debt