data quality assessment

Designs and implements quality assurance and quality control protocols, pipelines, and procedures that compute and apply objective, automated quality metrics and filters for datasets and embeddings. Builds rubrics and multi-criteria evaluation methods (including sample-level metrics, evidence-qualified evaluation, duplicate/leakage detection, distributional and SAR-artifact shift detection, and downstream-task utility assessment) and operationalizes aggregation of human feedback for reliable data quality assessment.

dataqualityassessment

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
2.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$203K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

How to Define the Quality of Data? A Feature-Based Literature Survey

Apr 02, 2025
MM
Markus Matoni
🏛️ Gesellschaft für wissenschaftliche Datenverarbeitung mbH | Philipps-Universität Marburg

Data quality definitions have long suffered from multidimensionality and conceptual inconsistency, necessitating systematic synthesis to establish a unified theoretical framework. This study introduces Feature-Oriented Domain Analysis (FODA) to data quality research for the first time, integrating a Systematic Literature Review (SLR) with quality dimension modeling to construct the first structured taxonomy encompassing mainstream definitions. We identify and clarify 12 core quality dimensions and their semantic relationships, proposing a novel four-level, feature-oriented taxonomy that significantly enhances definitional comparability and theoretical coherence. Our analysis reveals three critical research gaps: (1) lack of understanding of dynamic dimension evolution, (2) insufficient cross-domain semantic alignment, and (3) weak empirical validation. The resulting taxonomy provides a scalable, theoretically grounded foundation for data quality assessment, standardization, and tool development.

Classify existing data quality definitionsDefine multifaceted data quality dimensionsIdentify research gaps in data quality

Must-Read Papers

Most classic and influential ideas
View more

In highly regulated domains such as finance, data quality control (QC) is often fragmented into isolated preprocessing steps, undermining end-to-end trustworthy AI pipelines. To address this, we propose the first AI-driven DataOps framework that embeds QC as a system-level core component. Our framework deeply integrates rule-based engines, statistical analysis, and custom AI-powered anomaly detection across the entire data lifecycle—from ingestion and transformation to model deployment—enabling dynamic remediation, policy-configurable workflows, and end-to-end auditability. Technically, it unifies data profiling, stream processing, cloud-native storage interfaces, and a proprietary AI detection module. Evaluated in a real-world financial production environment, the framework achieves significantly improved anomaly recall, reduces manual intervention by 42%, and ensures audit completeness and full data traceability under high-throughput conditions, fully satisfying regulatory compliance requirements.

Enhances auditability and traceability in regulated data pipelinesIntegrates data quality control into continuous DataOps managementUnifies rule-based, statistical, and AI methods for anomaly detection

Development of Automated Data Quality Assessment and Evaluation Indices by Analytical Experience

Apr 03, 2025
YH
Yuka Haruki
🏛️ The University of Tokyo | Kyodo Printing Co., Ltd.

In data trading, expert-dependent Data Quality Assessment (DQA) impedes cross-organizational consensus, while practitioner experience heterogeneity exacerbates assessment bias. To address this, we conduct the first empirical study integrating eye-tracking with controlled comparative experiments to quantify how domain expertise influences perception and interpretation of quality metadata. Building on these findings, we propose a hierarchical, experience-adaptive DQA support paradigm and develop an automated tool for generating interpretable, multidimensional data quality metadata. The tool embeds explainable AI principles and integrates syntactic, semantic, and contextual quality indicators. Experimental evaluation demonstrates statistically significant reductions in DQA misclassification rates (p < 0.01) and improved usability for novice users. This work provides both theoretical foundations and empirical evidence for building trustworthy, transparent, and broadly deployable data quality infrastructure.

Automating data quality assessment to reduce expert dependencyComparing DQA performance between experienced and inexperienced usersDeveloping tools to improve data evaluation accuracy and consensus

This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.

data qualitydomain expertsno-code

Towards AI-Augmented Data Quality Management: From Data Quality for AI to AI for Data Quality Management

Jun 16, 2024
HC
Heidi Carolina Tamm
🏛️ Swedbank Group | University of Tartu

Data quality (DQ) rule definition in data warehouses remains heavily manual, resulting in low efficiency and high operational costs. Method: We systematically evaluated 151 industrial-grade DQ tools and conducted a comprehensive literature review to assess AI-enabled capabilities for automated DQ rule discovery in warehouse environments. Contribution/Results: Our analysis quantitatively reveals that only 10 tools exhibit preliminary AI-driven DQ rule detection capabilities—highlighting a significant gap in both industry practice and academic research. To address this, we propose the “AI for DQ Management” paradigm, shifting DQ governance from manual rule specification toward AI-autonomous rule discovery. We introduce a capability mapping matrix and a cross-platform functional comparison framework to precisely identify critical technical bottlenecks. This work provides empirically grounded guidance for organizational tool selection and outlines a research and development roadmap for next-generation, AI-native DQ governance systems.

Addressing market gap in AI-driven DQ toolsAssessing AI-augmented DQ rule detection capabilitiesAutomating data quality management in data warehouses

Improving ML Training Data with Gold-Standard Quality Metrics

Dec 23, 2025
LB
Leslie Barrett
🏛️ Bloomberg LP | Google

Manual annotation suffers from inconsistent quality and lacks systematic evaluation. Method: This paper proposes a consensus-based quality measurement framework grounded in multi-round annotation statistics, using dynamic decay of inter-annotator agreement variance as the core metric—established here as a “gold standard” for data quality. Recognizing annotators’ significant warm-up period but prohibitive cost of full-sample multiple annotation, we design a low-redundancy, high-efficiency progressive annotation protocol. The approach integrates statistical consistency analysis, variance convergence modeling, and label confidence estimation. Contribution/Results: Our paradigm substantially enhances data quality’s measurability and controllability: experiments show 3.2–7.8% accuracy gains across multiple NLP tasks and over 30% reduction in annotation redundancy.

Collecting high-quality training data without requiring multiple tags per itemEnhancing data quality through iterative tagging to reduce varianceEvaluating hand-tagged training data quality using statistical consistency metrics

Latest Papers

What's happening recently
View more

This study addresses the compliance challenges faced by data practitioners in machine learning systems under regulations such as the GDPR and the AI Act, particularly concerning data quality. Through semi-structured interviews with practitioners in the European Union, combined with thematic analysis of regulatory texts and engineering workflows, the research systematically uncovers a structural disconnect between regulation-driven data quality requirements and ML engineering practices. It identifies five core challenges: misalignment between legal principles and engineering implementation, fragmented data pipelines, lack of purpose-built compliance tools, ambiguous accountability, and reactive responses to audits. Building on these findings, the work proposes directions for designing compliance-oriented tooling, establishing effective governance mechanisms, and fostering cultural transformation to bridge the gap between regulatory mandates and practical ML development.

AI Actdata qualityGDPR

This study addresses the lack of systematic evaluation of data quality tools with respect to their measurement capabilities and integration with large language models (LLMs). It presents the first multidimensional assessment framework grounded in real-world enterprise use cases, systematically evaluating six prominent tools—including open-source solutions such as Great Expectations and Deequ, as well as commercial platforms like Informatica and Experian—across dimensions including rule definition, duplicate detection, metric aggregation, and uncertainty handling, along with their LLM integration mechanisms. The findings reveal that commercial tools offer more comprehensive functionality and初步 support for LLM-assisted rule generation, whereas open-source tools provide greater flexibility at the cost of higher implementation effort. Notably, none of the evaluated tools currently enable direct LLM-based data validation. This work provides empirical guidance for selecting data quality tools and advancing their integration with LLMs.

data qualitydata validationLLM integration

This work addresses the lack of a universal, tunable, and multi-scenario-compatible metric for data quality assessment, which hinders effective comparison of diverse data cleaning pipelines. To overcome this limitation, the authors propose TOMME—a general-purpose data quality measurement framework based on weighted errors—that extends traditional accuracy into a configurable, composite score. By producing a single quantitative metric, TOMME enables flexible adjustment of error weights according to specific use cases, thereby supporting both automated processing and optimization requirements. Experimental results demonstrate that TOMME exhibits strong adaptability, practicality, and comparability across a variety of scenarios, offering an efficient and unified solution for data quality evaluation and decision-making.

accuracydata cleaningdata quality

This work addresses the complex challenges of continuously monitoring item pool quality and health in large-scale AI-driven assessments. It proposes AQuAP, a dashboard system integrated with an item factory framework that leverages operational data analytics to support item generation and pool management. The system introduces novel metrics such as Effective Bank Size (EBS), which combines exposure rates and usage frequency to holistically evaluate the security, diversity, and efficiency of the item pool. By integrating psychometric indicators, exposure control algorithms, and advanced visualization techniques, AQuAP enables real-time monitoring of item pool vitality. The system has been successfully deployed in the Duolingo English Test, significantly enhancing the intelligence and responsiveness of item pool management.

AI-driven testingeducational assessmentitem bank health

Hot Scholars

GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
HD

Huiyu Duan

Shanghai Jiao Tong University
Multimedia Signal Processing
CX

Caiming Xiong

Salesforce Research
Machine LearningNLPComputer VisionMultimedia
JL

Junyang Lin

Qwen Team, Alibaba Group & Peking University
Natural Language ProcessingCross-Modal Representation LearningPretraining
PL

Peng Liang

School of Computer Science, Wuhan University
Software EngineeringSoftware ArchitectureEmpirical Software Engineering