Score
Manually examines data artifacts—labels, annotations, transformed records, and source mappings—by sampling records and comparing across sources to identify annotation inconsistencies, labeling discrepancies, and errors. Documents observed fidelity issues and concrete examples to inform relabeling, correction, or transformation fixes.
本文针对数据质量评估方法的文献空白,通过引入理论框架分类了四种主要方法和十五种具体手段,旨在为学术界提供理论基础,并帮助实践者做出明智决策。
A significant gap exists between theoretical definitions of data quality dimensions—such as accuracy, completeness, consistency, and timeliness—as stipulated in standards (e.g., ISO/IEC 25012) and the actual functionalities implemented in widely used data quality tools, with no systematic mapping analysis to date. Method: We conducted a systematic literature review and performed functional reverse engineering on seven mainstream open-source data quality tools. Contribution/Results: We present the first many-to-many mapping framework linking data quality dimensions to concrete tool capabilities, introducing a cross-dimensional, fine-grained functional categorization schema and generating a structured correspondence matrix covering all core dimensions. This work bridges the theory–practice divide, delivers an actionable guide for data quality assessment, and substantially enhances the scientific rigor and reusability of tool selection, functionality design, and standard implementation.
This study addresses the overreliance on inter-annotator agreement in current data annotation practices, which often overlooks annotation’s capacity to capture conceptual validity as a measurement act. Treating annotation as a measurement process, the work identifies five root causes of annotation issues—errors, ambiguity, impossibility, subjectivity, and annotator identity—and develops a measurement theory–based framework for diagnosing and improving annotation quality. Drawing on a synthesis of 132 literature sources and 10 semi-structured interviews, the research systematically defines target constructs, designs annotation instruments, implements labeling procedures, and evaluates both reliability and validity. The resulting framework equips annotation teams with evaluation methods that transcend mere agreement metrics, thereby substantially strengthening the foundational quality of AI training data.
Medical imaging datasets commonly suffer from label noise, shortcut learning, missing metadata, and challenges in retrospectively addressing newly discovered issues (e.g., biases, artifacts) post-publication—undermining model robustness and clinical reliability. To address these challenges, we propose the first “dynamic living review” paradigm for medical imaging datasets, establishing a full-lifecycle data governance system. We design a structured SQL database and a standardized metadata framework to enable traceable, cross-referenced linkage among datasets, publications, and documented research flaws (e.g., biases, annotation errors, shortcut effects). Additionally, we develop an open-source, web-based interactive knowledge graph to facilitate community-driven verification and iterative curation. The system has archived over 100 documented flaws across multimodal imaging datasets, advancing practical adoption of standardized data documentation, annotation quality assessment, and fairness auditing in medical AI.
This study addresses the challenges of quality assessment and contamination detection in large model training data, where absolute ground truth is often unavailable. We introduce the concept of "internal annotator dynamics," which analyzes a single annotator's self-reproducibility consistency on identical texts over time. Using the segmentation and causal-functional annotation of Sumerian mythology as an experimental testbed, we compare annotations produced by human experts and large language models (LLMs) across different time points. Our findings reveal that human annotators exhibit a distinctive pattern characterized by "stable boundaries but drifting labels." We demonstrate that this temporal signature can serve as an effective metric for data quality assessment and contamination detection, enabling the reliable identification of annotations that appear human-generated but are actually produced by LLMs.
Widespread label noise (10–25%) in NLP benchmark datasets leads to systematic underestimation of model performance, with many purported “LLM failures” attributable to annotation errors rather than model limitations. Method: We propose LLM-as-a-judge—a framework leveraging ensemble judgments from GPT-4, Claude, and Llama, combined with consistency voting and error-sensitivity analysis to automatically detect mislabeled instances; we further apply label smoothing and confident learning for robust label recalibration. Contribution/Results: Comprehensive evaluation across the TRUE benchmark suite reveals substantial disparities in quality and efficiency among expert, crowdsourced, and LLM-generated annotations. After correction, state-of-the-art models achieve average accuracy gains of 3.2–7.8 percentage points. This work provides the first empirical evidence of systematic label-noise interference in LLM evaluation and introduces a scalable, collaborative adjudication paradigm that reframes data correction as model performance recalibration.
NLP data quality assessment has long relied on inter-annotator agreement, overlooking intra-annotator consistency—the temporal stability of individual annotators’ judgments. This neglect challenges the implicit “gold label as ground truth” assumption. Method: We conduct exploratory repeated annotation experiments across major NLP datasets and quantify intra-annotator agreement using Cohen’s and Fleiss’ Kappa, complemented by qualitative perceptual analysis. Contribution/Results: We demonstrate that mainstream NLP datasets routinely omit intra-annotator consistency reporting; moreover, individual annotators exhibit significant temporal variability in labeling identical texts. We identify and disentangle the dual influence of textual ambiguity and subjectivity on annotation stability. Our work establishes intra-annotator agreement as a foundational data quality metric, providing both a methodological framework and concrete guidelines for constructing more robust, reproducible NLP datasets.
Large language models are prone to generating hallucinations or dubious citations in academic writing, undermining research credibility. This study presents the first systematic evaluation and comparison of mainstream citation verification tools—CheckIfExist, HalluCiteChecker, Hallucinator, HalRef, and RefChecker—on real-world academic documents. The analysis reveals significant limitations in current approaches, particularly concerning citation extraction accuracy, breadth of database coverage, and consistency in verification. While these tools can offer preliminary alerts for potentially fabricated references, their overall effectiveness remains constrained. This work provides an empirical foundation and clear directions for improving the verification of citation authenticity in scholarly communication.
This study addresses the absence of a universally accepted and empirically validated definition of descriptive data quality in the cultural heritage domain. It proposes, for the first time, a domain-specific data quality dimension framework developed through a systematic empirical evaluation that integrates literature review, dimensional analysis, and problem annotation on real-world datasets. By bridging the gap between data quality theory and cultural heritage practice, this work delivers an actionable, contextually tailored, and empirically grounded definition of data quality. The resulting framework provides clear guidance for digital initiatives in the field, supporting more effective data curation, interoperability, and long-term preservation efforts within cultural heritage institutions.
Real-world data often contain label noise that severely degrades model performance. This work proposes Relabeler, a novel framework that, for the first time, jointly models both local and global relationships among data instances within a unified architecture to enable end-to-end label correction. By leveraging input features and observed labels in concert, Relabeler estimates the most probable clean labels through a data-centric learning paradigm that integrates relational modeling, probabilistic label inference, and end-to-end optimization. Extensive experiments demonstrate that the method consistently outperforms existing approaches across diverse datasets, noise types, and noise rates, achieving up to a 58% improvement in label correction accuracy and a 6% gain in downstream task performance.
This study addresses the deep semantic inconsistencies between clinical notes and structured tables in electronic health records (EHRs), which existing methods fail to capture due to their reliance on superficial matching rather than clinical reasoning, event relationships, and temporal dynamics. To tackle this, the authors introduce EHR-ReasonCon, a reasoning-intensive consistency verification benchmark built on MIMIC-III, featuring the first expert-guided, inference-level annotation protocol and a dedicated table exploration tool. They further propose EHR-Inspector, an LLM-driven framework that integrates clinical text segmentation, anchor entity and temporal reference extraction, structured querying, and an LLM-as-a-judge evaluation mechanism. Experiments demonstrate consistent and significant improvements over baselines across multiple model backbones, with robust performance under both stringent and lenient expert evaluations. Ablation studies confirm the efficacy of the design and highlight nuanced discrepancies with human judgments.
研究通过审计117个开源科研软件项目的多个元数据表面,发现大部分项目存在核心字段冲突,影响软件引用一致性。