Development of Automated Data Quality Assessment and Evaluation Indices by Analytical Experience

📅 2025-04-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
In data trading, expert-dependent Data Quality Assessment (DQA) impedes cross-organizational consensus, while practitioner experience heterogeneity exacerbates assessment bias. To address this, we conduct the first empirical study integrating eye-tracking with controlled comparative experiments to quantify how domain expertise influences perception and interpretation of quality metadata. Building on these findings, we propose a hierarchical, experience-adaptive DQA support paradigm and develop an automated tool for generating interpretable, multidimensional data quality metadata. The tool embeds explainable AI principles and integrates syntactic, semantic, and contextual quality indicators. Experimental evaluation demonstrates statistically significant reductions in DQA misclassification rates (p < 0.01) and improved usability for novice users. This work provides both theoretical foundations and empirical evidence for building trustworthy, transparent, and broadly deployable data quality infrastructure.

Technology Category

Data Mining & Knowledge Management: Representing, Reasoning, and Using Provenance, TrustKnowledge Representation and Reasoning: Qualitative ReasoningSearch and Optimization: Metareasoning and Metaheuristics

Application Category

Economics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsWeb Mining and Content Analysis: Web data quality in the era of algorithmically-generated contentSearch and Retrieval-Augmented AI: Web evaluation methodologies and metrics
📝 Abstract
The societal need to leverage third-party data has driven the data-distribution market and increased the importance of data quality assessment (DQA) in data transactions between organizations. However, DQA requires expert knowledge of raw data and related data attributes, which hinders consensus-building in data purchasing. This study focused on the differences in DQAs between experienced and inexperienced data handlers. We performed two experiments: The first was a questionnaire survey involving 41 participants with varying levels of data-handling experience, who evaluated 12 data samples using 10 predefined indices with and without quality metadata generated by the automated tool. The second was an eye-tracking experiment to reveal the viewing behavior of participants during data evaluation. It was revealed that using quality metadata generated by the automated tool can reduce misrecognition in DQA. While experienced data handlers rated the quality metadata highly, semi-experienced users gave it the lowest ratings. This study contributes to enhancing data understanding within organizations and promoting the distribution of valuable data by proposing an automated tool to support DQAs.
Problem

Research questions and friction points this paper is trying to address.

Automating data quality assessment to reduce expert dependency
Comparing DQA performance between experienced and inexperienced users
Developing tools to improve data evaluation accuracy and consensus
Innovation

Methods, ideas, or system contributions that make the work stand out.

Automated tool generates quality metadata
Questionnaire and eye-tracking experiments conducted
Reduces misrecognition in data quality assessment
💼 Related Jobs
No related jobs found.
Y
Yuka Haruki
School of Engineering, The University of Tokyo.
K
Kei Kato
Kyodo Printing Co., Ltd..
Y
Yuki Enami
Kyodo Printing Co., Ltd..
H
Hiroaki Takeuchi
Kyodo Printing Co., Ltd..
D
Daiki Kazuno
Kyodo Printing Co., Ltd..
K
Kotaro Yamada
Kyodo Printing Co., Ltd..
T
Teruaki Hayashi
School of Engineering, The University of Tokyo.