Score
Designs and builds systems, frameworks, and operational processes that detect, quantify, track, and correct data errors—implementing error correction and mitigation techniques (including forward error correction), managed correction loops, and error-tracking mechanisms—and that enforce data quality checks, validation, control, and monitoring to maintain integrity. Defines and applies data quality metrics, assessment, assurance, and auditing practices to measure the effectiveness of those controls and to drive continuous improvement of data quality.
A significant gap exists between theoretical definitions of data quality dimensions—such as accuracy, completeness, consistency, and timeliness—as stipulated in standards (e.g., ISO/IEC 25012) and the actual functionalities implemented in widely used data quality tools, with no systematic mapping analysis to date. Method: We conducted a systematic literature review and performed functional reverse engineering on seven mainstream open-source data quality tools. Contribution/Results: We present the first many-to-many mapping framework linking data quality dimensions to concrete tool capabilities, introducing a cross-dimensional, fine-grained functional categorization schema and generating a structured correspondence matrix covering all core dimensions. This work bridges the theory–practice divide, delivers an actionable guide for data quality assessment, and substantially enhances the scientific rigor and reusability of tool selection, functionality design, and standard implementation.
Data quality definitions have long suffered from multidimensionality and conceptual inconsistency, necessitating systematic synthesis to establish a unified theoretical framework. This study introduces Feature-Oriented Domain Analysis (FODA) to data quality research for the first time, integrating a Systematic Literature Review (SLR) with quality dimension modeling to construct the first structured taxonomy encompassing mainstream definitions. We identify and clarify 12 core quality dimensions and their semantic relationships, proposing a novel four-level, feature-oriented taxonomy that significantly enhances definitional comparability and theoretical coherence. Our analysis reveals three critical research gaps: (1) lack of understanding of dynamic dimension evolution, (2) insufficient cross-domain semantic alignment, and (3) weak empirical validation. The resulting taxonomy provides a scalable, theoretically grounded foundation for data quality assessment, standardization, and tool development.
In highly regulated domains such as finance, data quality control (QC) is often fragmented into isolated preprocessing steps, undermining end-to-end trustworthy AI pipelines. To address this, we propose the first AI-driven DataOps framework that embeds QC as a system-level core component. Our framework deeply integrates rule-based engines, statistical analysis, and custom AI-powered anomaly detection across the entire data lifecycle—from ingestion and transformation to model deployment—enabling dynamic remediation, policy-configurable workflows, and end-to-end auditability. Technically, it unifies data profiling, stream processing, cloud-native storage interfaces, and a proprietary AI detection module. Evaluated in a real-world financial production environment, the framework achieves significantly improved anomaly recall, reduces manual intervention by 42%, and ensures audit completeness and full data traceability under high-throughput conditions, fully satisfying regulatory compliance requirements.
The core challenge in data quality monitoring lies in error provenance—specifically, identifying the underlying mechanisms that generate errors—a problem largely overlooked by existing work, which seldom models such mechanisms explicitly. This paper focuses on errors arising from intrinsic dependencies within data and proposes MechDetect, the first method to systematically extend missing-data mechanism detection to diverse error types—including outliers, inconsistencies, and format violations. Leveraging joint statistical modeling and supervised learning, MechDetect simultaneously models tabular data and their error masks to automatically determine whether observed errors stem from inherent characteristics of the original data. Extensive experiments across multiple benchmark datasets demonstrate that MechDetect significantly outperforms state-of-the-art baselines in accurately diagnosing error-generation mechanisms. By providing mechanistic interpretability, it establishes a theoretical foundation and practical framework for explainable data repair.
Ad hoc SQL development lacks engineering rigor, leading to data silos, logical redundancy, and ineffective data governance. Method: This paper proposes a DataOps-driven CI/CD framework for analytical SQL warehouses, featuring a novel five-stage automated pipeline—Lint, Optimize, Parse, Validate, Observe—that embeds quality assurance and enables end-to-end lifecycle governance. Contribution/Results: We introduce the DataOps Controls Scorecard and a requirements traceability matrix, explicitly mapping 12 governance criteria to CI/CD stages to ensure control completeness and scalability. The framework integrates Agile, Lean, and DevOps principles with static analysis, syntactic parsing, optimization recommendations, validation testing, and observability. Empirical evaluation demonstrates significant improvements in data quality, development transparency, and cross-functional collaboration, providing a sustainable, production-ready pathway for large-scale analytical systems.
This study addresses the compliance challenges faced by data practitioners in machine learning systems under regulations such as the GDPR and the AI Act, particularly concerning data quality. Through semi-structured interviews with practitioners in the European Union, combined with thematic analysis of regulatory texts and engineering workflows, the research systematically uncovers a structural disconnect between regulation-driven data quality requirements and ML engineering practices. It identifies five core challenges: misalignment between legal principles and engineering implementation, fragmented data pipelines, lack of purpose-built compliance tools, ambiguous accountability, and reactive responses to audits. Building on these findings, the work proposes directions for designing compliance-oriented tooling, establishing effective governance mechanisms, and fostering cultural transformation to bridge the gap between regulatory mandates and practical ML development.
This study addresses the pervasive issue of data errors in real-world databases—such as missing values, redundancy, statistical biases, and outliers—which significantly degrade downstream analytical and machine learning performance. Recognizing that existing taxonomies are incomplete and terminology inconsistent, this work presents the first unified framework that integrates traditional data errors with statistically oriented inaccuracies critical in the AI era. It proposes a non-overlapping tripartite classification structure—comprising missing, erroneous, and redundant data—and systematically constructs a comprehensive catalog of 35 distinct error types. Through formal definitions, illustrative examples, and a thorough literature review, the paper establishes standardized terminology and precise characterizations, thereby offering a clear, rigorous theoretical foundation and practical toolkit for data quality assessment and cleaning.
This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.
This study addresses the lack of systematic evaluation of data quality tools with respect to their measurement capabilities and integration with large language models (LLMs). It presents the first multidimensional assessment framework grounded in real-world enterprise use cases, systematically evaluating six prominent tools—including open-source solutions such as Great Expectations and Deequ, as well as commercial platforms like Informatica and Experian—across dimensions including rule definition, duplicate detection, metric aggregation, and uncertainty handling, along with their LLM integration mechanisms. The findings reveal that commercial tools offer more comprehensive functionality and初步 support for LLM-assisted rule generation, whereas open-source tools provide greater flexibility at the cost of higher implementation effort. Notably, none of the evaluated tools currently enable direct LLM-based data validation. This work provides empirical guidance for selecting data quality tools and advancing their integration with LLMs.
This work addresses the lack of a universal, tunable, and multi-scenario-compatible metric for data quality assessment, which hinders effective comparison of diverse data cleaning pipelines. To overcome this limitation, the authors propose TOMME—a general-purpose data quality measurement framework based on weighted errors—that extends traditional accuracy into a configurable, composite score. By producing a single quantitative metric, TOMME enables flexible adjustment of error weights according to specific use cases, thereby supporting both automated processing and optimization requirements. Experimental results demonstrate that TOMME exhibits strong adaptability, practicality, and comparability across a variety of scenarios, offering an efficient and unified solution for data quality evaluation and decision-making.
This study addresses the challenge of unreliable AI models and diminished clinical trust stemming from opaque data quality reporting in the secondary use of electronic health records (EHRs). To this end, the authors propose the first comprehensive framework for transparent data quality reporting across the entire EHR lifecycle. The framework innovatively distinguishes between data producers and consumers, explicitly defines five critical phases, and maps established data quality dimensions to specific workflow stages. Through iterative stakeholder and process analysis, a structured reporting mechanism is developed and validated on real-world datasets, demonstrating its ability to effectively trace the origins of data quality issues. The approach significantly enhances data interpretability, fitness-for-use, and governance efficacy, thereby providing a robust foundation for trustworthy AI development and clinical research.