perform qualitative inspection

Manually examines data artifacts—labels, annotations, transformed records, and source mappings—by sampling records and comparing across sources to identify annotation inconsistencies, labeling discrepancies, and errors. Documents observed fidelity issues and concrete examples to inform relabeling, correction, or transformation fixes.

performqualitativeinspection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.42
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$212K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Unfolding Data Quality Dimensions in Practice: A Survey

Jul 23, 2025
VP
Vasileios Papastergios
🏛️ Aristotle University | Hasso Plattner Institute | University of Potsdam

A significant gap exists between theoretical definitions of data quality dimensions—such as accuracy, completeness, consistency, and timeliness—as stipulated in standards (e.g., ISO/IEC 25012) and the actual functionalities implemented in widely used data quality tools, with no systematic mapping analysis to date. Method: We conducted a systematic literature review and performed functional reverse engineering on seven mainstream open-source data quality tools. Contribution/Results: We present the first many-to-many mapping framework linking data quality dimensions to concrete tool capabilities, introducing a cross-dimensional, fine-grained functional categorization schema and generating a structured correspondence matrix covering all core dimensions. This work bridges the theory–practice divide, delivers an actionable guide for data quality assessment, and substantially enhances the scientific rigor and reusability of tool selection, functionality design, and standard implementation.

Bridging gap between data quality theory and practiceMapping tool functionalities to data quality dimensionsProviding unified view on fragmented quality checks

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the overreliance on inter-annotator agreement in current data annotation practices, which often overlooks annotation’s capacity to capture conceptual validity as a measurement act. Treating annotation as a measurement process, the work identifies five root causes of annotation issues—errors, ambiguity, impossibility, subjectivity, and annotator identity—and develops a measurement theory–based framework for diagnosing and improving annotation quality. Drawing on a synthesis of 132 literature sources and 10 semi-structured interviews, the research systematically defines target constructs, designs annotation instruments, implements labeling procedures, and evaluates both reliability and validity. The resulting framework equips annotation teams with evaluation methods that transcend mere agreement metrics, thereby substantially strengthening the foundational quality of AI training data.

annotation qualitydata annotationmeasurement

In the Picture: Medical Imaging Datasets, Artifacts, and their Living Review

Jan 18, 2025
AJ
Amelia Jim'enez-S'anchez
🏛️ IT University of Copenhagen | University of Copenhagen | Radboud University Medical Center | Universitat de Barcelona | Technical University of Denmark | CONICET | University of Buenos Aires | Emory University | Stanford University | University of Groningen | Aarhus University | The Hebrew University of Jerusalem | University of Southern Denmark | Lunit | Cerebriu A/S | Federal University of Espírito Santo | German Cancer Research Center | Heidelberg University | University of Bern | Plain Medical | Oxford

Medical imaging datasets commonly suffer from label noise, shortcut learning, missing metadata, and challenges in retrospectively addressing newly discovered issues (e.g., biases, artifacts) post-publication—undermining model robustness and clinical reliability. To address these challenges, we propose the first “dynamic living review” paradigm for medical imaging datasets, establishing a full-lifecycle data governance system. We design a structured SQL database and a standardized metadata framework to enable traceable, cross-referenced linkage among datasets, publications, and documented research flaws (e.g., biases, annotation errors, shortcut effects). Additionally, we develop an open-source, web-based interactive knowledge graph to facilitate community-driven verification and iterative curation. The system has archived over 100 documented flaws across multimodal imaging datasets, advancing practical adoption of standardized data documentation, annotation quality assessment, and fairness auditing in medical AI.

Algorithm PerformanceDataset QualityMedical Image Analysis

This study addresses the challenges of quality assessment and contamination detection in large model training data, where absolute ground truth is often unavailable. We introduce the concept of "internal annotator dynamics," which analyzes a single annotator's self-reproducibility consistency on identical texts over time. Using the segmentation and causal-functional annotation of Sumerian mythology as an experimental testbed, we compare annotations produced by human experts and large language models (LLMs) across different time points. Our findings reveal that human annotators exhibit a distinctive pattern characterized by "stable boundaries but drifting labels." We demonstrate that this temporal signature can serve as an effective metric for data quality assessment and contamination detection, enabling the reliable identification of annotations that appear human-generated but are actually produced by LLMs.

annotation consistencydata qualityintra-annotator dynamics

Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performance

Oct 24, 2024
ON
Omer Nahum
🏛️ Technion - Institute of Technology | Google Research

Widespread label noise (10–25%) in NLP benchmark datasets leads to systematic underestimation of model performance, with many purported “LLM failures” attributable to annotation errors rather than model limitations. Method: We propose LLM-as-a-judge—a framework leveraging ensemble judgments from GPT-4, Claude, and Llama, combined with consistency voting and error-sensitivity analysis to automatically detect mislabeled instances; we further apply label smoothing and confident learning for robust label recalibration. Contribution/Results: Comprehensive evaluation across the TRUE benchmark suite reveals substantial disparities in quality and efficiency among expert, crowdsourced, and LLM-generated annotations. After correction, state-of-the-art models achieve average accuracy gains of 3.2–7.8 percentage points. This work provides the first empirical evidence of systematic label-noise interference in LLM evaluation and introduces a scalable, collaborative adjudication paradigm that reframes data correction as model performance recalibration.

Comparing annotation quality from experts, crowdsourcing, and LLMsDetecting label errors in NLP benchmark datasetsMitigating mislabeled data effects on model performance

Consistency is Key: Disentangling Label Variation in Natural Language Processing with Intra-Annotator Agreement

Jan 25, 2023
GA
Gavin Abercrombie
🏛️ Heriot-Watt University | Alana AI | Bocconi University

NLP data quality assessment has long relied on inter-annotator agreement, overlooking intra-annotator consistency—the temporal stability of individual annotators’ judgments. This neglect challenges the implicit “gold label as ground truth” assumption. Method: We conduct exploratory repeated annotation experiments across major NLP datasets and quantify intra-annotator agreement using Cohen’s and Fleiss’ Kappa, complemented by qualitative perceptual analysis. Contribution/Results: We demonstrate that mainstream NLP datasets routinely omit intra-annotator consistency reporting; moreover, individual annotators exhibit significant temporal variability in labeling identical texts. We identify and disentangle the dual influence of textual ambiguity and subjectivity on annotation stability. Our work establishes intra-annotator agreement as a foundational data quality metric, providing both a methodological framework and concrete guidelines for constructing more robust, reproducible NLP datasets.

Assessing annotator inconsistency across multiple NLP classification tasksInvestigating reasons for annotator disagreement through quality control measuresMeasuring intra-annotator agreement for label stability in NLP tasks

Latest Papers

What's happening recently
View more

Large language models are prone to generating hallucinations or dubious citations in academic writing, undermining research credibility. This study presents the first systematic evaluation and comparison of mainstream citation verification tools—CheckIfExist, HalluCiteChecker, Hallucinator, HalRef, and RefChecker—on real-world academic documents. The analysis reveals significant limitations in current approaches, particularly concerning citation extraction accuracy, breadth of database coverage, and consistency in verification. While these tools can offer preliminary alerts for potentially fabricated references, their overall effectiveness remains constrained. This work provides an empirical foundation and clear directions for improving the verification of citation authenticity in scholarly communication.

academic writinghallucinated citationsreference reliability

This study addresses the absence of a universally accepted and empirically validated definition of descriptive data quality in the cultural heritage domain. It proposes, for the first time, a domain-specific data quality dimension framework developed through a systematic empirical evaluation that integrates literature review, dimensional analysis, and problem annotation on real-world datasets. By bridging the gap between data quality theory and cultural heritage practice, this work delivers an actionable, contextually tailored, and empirically grounded definition of data quality. The resulting framework provides clear guidance for digital initiatives in the field, supporting more effective data curation, interoperability, and long-term preservation efforts within cultural heritage institutions.

cultural heritagedata qualitydescriptive information

Real-world data often contain label noise that severely degrades model performance. This work proposes Relabeler, a novel framework that, for the first time, jointly models both local and global relationships among data instances within a unified architecture to enable end-to-end label correction. By leveraging input features and observed labels in concert, Relabeler estimates the most probable clean labels through a data-centric learning paradigm that integrates relational modeling, probabilistic label inference, and end-to-end optimization. Extensive experiments demonstrate that the method consistently outperforms existing approaches across diverse datasets, noise types, and noise rates, achieving up to a 58% improvement in label correction accuracy and a 6% gain in downstream task performance.

data qualitylabel correctionlabel corruption

This study addresses the deep semantic inconsistencies between clinical notes and structured tables in electronic health records (EHRs), which existing methods fail to capture due to their reliance on superficial matching rather than clinical reasoning, event relationships, and temporal dynamics. To tackle this, the authors introduce EHR-ReasonCon, a reasoning-intensive consistency verification benchmark built on MIMIC-III, featuring the first expert-guided, inference-level annotation protocol and a dedicated table exploration tool. They further propose EHR-Inspector, an LLM-driven framework that integrates clinical text segmentation, anchor entity and temporal reference extraction, structured querying, and an LLM-as-a-judge evaluation mechanism. Experiments demonstrate consistent and significant improvements over baselines across multiple model backbones, with robust performance under both stringent and lenient expert evaluations. Ablation studies confirm the efficacy of the design and highlight nuanced discrepancies with human judgments.

clinical notesdata inconsistencyEHR consistency

Hot Scholars

CJ

Cho-Jui Hsieh

University of California, Los Angeles
Machine LearningOptimization
CD

Chao Dong

Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences
image restorationincluding super-resolutiondenoisingetc.
JL

Juntao Li

Soochow University
Language ModelsText Generation
HW

Haitian Wang

University of Western Australia
3D point cloudComputer visionMachine leaningIoT
QQ

Quantong Qiu

Soochow University
LLMSparse AttentionKV Cache