data quality engineering

Designs and builds systems, frameworks, and operational processes that detect, quantify, track, and correct data errors—implementing error correction and mitigation techniques (including forward error correction), managed correction loops, and error-tracking mechanisms—and that enforce data quality checks, validation, control, and monitoring to maintain integrity. Defines and applies data quality metrics, assessment, assurance, and auditing practices to measure the effectiveness of those controls and to drive continuous improvement of data quality.

dataqualityengineering

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.88
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$195K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

How to Define the Quality of Data? A Feature-Based Literature Survey

Apr 02, 2025
MM
Markus Matoni
🏛️ Gesellschaft für wissenschaftliche Datenverarbeitung mbH | Philipps-Universität Marburg

Data quality definitions have long suffered from multidimensionality and conceptual inconsistency, necessitating systematic synthesis to establish a unified theoretical framework. This study introduces Feature-Oriented Domain Analysis (FODA) to data quality research for the first time, integrating a Systematic Literature Review (SLR) with quality dimension modeling to construct the first structured taxonomy encompassing mainstream definitions. We identify and clarify 12 core quality dimensions and their semantic relationships, proposing a novel four-level, feature-oriented taxonomy that significantly enhances definitional comparability and theoretical coherence. Our analysis reveals three critical research gaps: (1) lack of understanding of dynamic dimension evolution, (2) insufficient cross-domain semantic alignment, and (3) weak empirical validation. The resulting taxonomy provides a scalable, theoretically grounded foundation for data quality assessment, standardization, and tool development.

Classify existing data quality definitionsDefine multifaceted data quality dimensionsIdentify research gaps in data quality

Must-Read Papers

Most classic and influential ideas
View more

In highly regulated domains such as finance, data quality control (QC) is often fragmented into isolated preprocessing steps, undermining end-to-end trustworthy AI pipelines. To address this, we propose the first AI-driven DataOps framework that embeds QC as a system-level core component. Our framework deeply integrates rule-based engines, statistical analysis, and custom AI-powered anomaly detection across the entire data lifecycle—from ingestion and transformation to model deployment—enabling dynamic remediation, policy-configurable workflows, and end-to-end auditability. Technically, it unifies data profiling, stream processing, cloud-native storage interfaces, and a proprietary AI detection module. Evaluated in a real-world financial production environment, the framework achieves significantly improved anomaly recall, reduces manual intervention by 42%, and ensures audit completeness and full data traceability under high-throughput conditions, fully satisfying regulatory compliance requirements.

Enhances auditability and traceability in regulated data pipelinesIntegrates data quality control into continuous DataOps managementUnifies rule-based, statistical, and AI methods for anomaly detection

MechDetect: Detecting Data-Dependent Errors

Dec 03, 2025
PJ
Philipp Jung
🏛️ Berlin University of Applied Sciences and Technology

The core challenge in data quality monitoring lies in error provenance—specifically, identifying the underlying mechanisms that generate errors—a problem largely overlooked by existing work, which seldom models such mechanisms explicitly. This paper focuses on errors arising from intrinsic dependencies within data and proposes MechDetect, the first method to systematically extend missing-data mechanism detection to diverse error types—including outliers, inconsistencies, and format violations. Leveraging joint statistical modeling and supervised learning, MechDetect simultaneously models tabular data and their error masks to automatically determine whether observed errors stem from inherent characteristics of the original data. Extensive experiments across multiple benchmark datasets demonstrate that MechDetect significantly outperforms state-of-the-art baselines in accurately diagnosing error-generation mechanisms. By providing mechanistic interpretability, it establishes a theoretical foundation and practical framework for explainable data repair.

Detect data-dependent error generation mechanismsEstimate error dependency using machine learning modelsExtend missing value analysis to other error types

DataOps-driven CI/CD for analytics repositories

Nov 15, 2025
DV
Dmytro Valiaiev
🏛️ University of Arkansas Little Rock

Ad hoc SQL development lacks engineering rigor, leading to data silos, logical redundancy, and ineffective data governance. Method: This paper proposes a DataOps-driven CI/CD framework for analytical SQL warehouses, featuring a novel five-stage automated pipeline—Lint, Optimize, Parse, Validate, Observe—that embeds quality assurance and enables end-to-end lifecycle governance. Contribution/Results: We introduce the DataOps Controls Scorecard and a requirements traceability matrix, explicitly mapping 12 governance criteria to CI/CD stages to ensure control completeness and scalability. The framework integrates Agile, Lean, and DevOps principles with static analysis, syntactic parsing, optimization recommendations, validation testing, and observability. Empirical evaluation demonstrates significant improvements in data quality, development transparency, and cross-functional collaboration, providing a sustainable, production-ready pathway for large-scale analytical systems.

Addressing ad-hoc SQL development lacking software engineering rigorProviding standardized DataOps framework for analytics pipeline managementSolving data governance challenges and validation impossibility in analytics

This study addresses the compliance challenges faced by data practitioners in machine learning systems under regulations such as the GDPR and the AI Act, particularly concerning data quality. Through semi-structured interviews with practitioners in the European Union, combined with thematic analysis of regulatory texts and engineering workflows, the research systematically uncovers a structural disconnect between regulation-driven data quality requirements and ML engineering practices. It identifies five core challenges: misalignment between legal principles and engineering implementation, fragmented data pipelines, lack of purpose-built compliance tools, ambiguous accountability, and reactive responses to audits. Building on these findings, the work proposes directions for designing compliance-oriented tooling, establishing effective governance mechanisms, and fostering cultural transformation to bridge the gap between regulatory mandates and practical ML development.

AI Actdata qualityGDPR

This study addresses the pervasive issue of data errors in real-world databases—such as missing values, redundancy, statistical biases, and outliers—which significantly degrade downstream analytical and machine learning performance. Recognizing that existing taxonomies are incomplete and terminology inconsistent, this work presents the first unified framework that integrates traditional data errors with statistically oriented inaccuracies critical in the AI era. It proposes a non-overlapping tripartite classification structure—comprising missing, erroneous, and redundant data—and systematically constructs a comprehensive catalog of 35 distinct error types. Through formal definitions, illustrative examples, and a thorough literature review, the paper establishes standardized terminology and precise characterizations, thereby offering a clear, rigorous theoretical foundation and practical toolkit for data quality assessment and cleaning.

data errorsdata qualityerror taxonomy

Latest Papers

What's happening recently
View more

This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.

data qualitydomain expertsno-code

This study addresses the lack of systematic evaluation of data quality tools with respect to their measurement capabilities and integration with large language models (LLMs). It presents the first multidimensional assessment framework grounded in real-world enterprise use cases, systematically evaluating six prominent tools—including open-source solutions such as Great Expectations and Deequ, as well as commercial platforms like Informatica and Experian—across dimensions including rule definition, duplicate detection, metric aggregation, and uncertainty handling, along with their LLM integration mechanisms. The findings reveal that commercial tools offer more comprehensive functionality and初步 support for LLM-assisted rule generation, whereas open-source tools provide greater flexibility at the cost of higher implementation effort. Notably, none of the evaluated tools currently enable direct LLM-based data validation. This work provides empirical guidance for selecting data quality tools and advancing their integration with LLMs.

data qualitydata validationLLM integration

This work addresses the lack of a universal, tunable, and multi-scenario-compatible metric for data quality assessment, which hinders effective comparison of diverse data cleaning pipelines. To overcome this limitation, the authors propose TOMME—a general-purpose data quality measurement framework based on weighted errors—that extends traditional accuracy into a configurable, composite score. By producing a single quantitative metric, TOMME enables flexible adjustment of error weights according to specific use cases, thereby supporting both automated processing and optimization requirements. Experimental results demonstrate that TOMME exhibits strong adaptability, practicality, and comparability across a variety of scenarios, offering an efficient and unified solution for data quality evaluation and decision-making.

accuracydata cleaningdata quality

This study addresses the challenge of unreliable AI models and diminished clinical trust stemming from opaque data quality reporting in the secondary use of electronic health records (EHRs). To this end, the authors propose the first comprehensive framework for transparent data quality reporting across the entire EHR lifecycle. The framework innovatively distinguishes between data producers and consumers, explicitly defines five critical phases, and maps established data quality dimensions to specific workflow stages. Through iterative stakeholder and process analysis, a structured reporting mechanism is developed and validated on real-world datasets, demonstrating its ability to effectively trace the origins of data quality issues. The approach significantly enhances data interpretability, fitness-for-use, and governance efficacy, thereby providing a robust foundation for trustworthy AI development and clinical research.

clinical AIdata lifecycledata quality

Hot Scholars

GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
HD

Huiyu Duan

Shanghai Jiao Tong University
Multimedia Signal Processing
LB

Lei Bai

Shanghai AI Laboratory
Foundation ModelScience IntelligenceMulti-Agent SystemAutonomous Discovery
JY

Junchi Yan

FIAPR & ICML Board Member, SJTU (2018-), SII (2024-), AWS (2019-2022), IBM (2011-2018)
Computational IntelligenceAI4ScienceMachine LearningAutonomous Driving