Score
Designs, builds, and validates labeled datasets and the processes that produce them by creating labeling schemas and guidelines, selecting or building annotation tools, training and coordinating annotators, and performing quality control and inter-annotator agreement analysis. Produces reliable ground-truth annotations (labels) and annotation pipelines needed to train, evaluate, or analyze data-driven systems.
High annotation costs and prolonged turnaround times plague NLP development, necessitating efficient and reliable data labeling paradigms. This paper proposes an LLM-powered Human-in-the-Loop (HITL) hybrid annotation framework that systematically integrates synthetic data generation, active learning, and human-AI collaboration, augmented with built-in mechanisms for annotation quality assessment, annotator management, and cost-benefit analysis. Unlike prior work—largely theoretical or narrowly scoped—this study introduces the first deployable, plug-and-play industrial-grade annotation methodology, bridging the critical gap between methodological research and real-world engineering practice. Empirical validation across multiple production NLP projects demonstrates that the framework consistently reduces annotation costs and cycle time by 30–50%, while maintaining label quality within required thresholds.
This study addresses the overreliance on inter-annotator agreement in current data annotation practices, which often overlooks annotation’s capacity to capture conceptual validity as a measurement act. Treating annotation as a measurement process, the work identifies five root causes of annotation issues—errors, ambiguity, impossibility, subjectivity, and annotator identity—and develops a measurement theory–based framework for diagnosing and improving annotation quality. Drawing on a synthesis of 132 literature sources and 10 semi-structured interviews, the research systematically defines target constructs, designs annotation instruments, implements labeling procedures, and evaluates both reliability and validity. The resulting framework equips annotation teams with evaluation methods that transcend mere agreement metrics, thereby substantially strengthening the foundational quality of AI training data.
This study addresses the widespread neglect in machine learning research of when validation occurs during data annotation—a critical factor influencing both label quality and cost—despite overreliance on post-hoc quality control. Drawing inspiration from the “shift-left” principle in software engineering, this work proposes a tripartite classification of quality checkpoints across early, intermediate, and late stages of the annotation pipeline and introduces a parameterized error propagation model that, for the first time, treats validation timing as a quantifiable design variable. Through error propagation modeling, process decomposition, and literature analysis, the authors find that only 4% of recent studies report validation timing. Their analysis demonstrates that early-stage quality checks can reduce error correction costs by up to two orders of magnitude. The paper calls for standardized reporting of timing configurations, platform support for tunable timing parameters, and empirical studies on stage-specific detection rates.
Scientific process descriptions are often embedded in unstructured text, hindering reproducibility, comparison, and automation. To address this challenge, this work presents the first cross-disciplinary, expert-driven repository of structured scientific process schemas, encompassing 16 expert-annotated patterns across five domains. Through a human-in-the-loop workflow, candidate schemas generated by large language models were iteratively refined via domain expert feedback, yielding reusable fields such as inputs, outputs, steps, and parameters. The resulting schemas are formalized in both JSON Schema and SHACL formats and accompanied by an integrated toolchain. The project also releases a comprehensive dataset—including schemas, intermediate artifacts, review records, and analysis scripts—to support knowledge graph construction, semantic publishing, and cross-study comparison.
In autonomous driving AI perception system (AIePS) development, annotation quality critically impacts model safety and reliability; however, empirical understanding of how annotation errors originate and propagate across multi-organizational automotive supply chains remains lacking. This study addresses this gap by adopting dual perspectives—annotation lifecycle and supply chain—through semi-structured interviews with 20 domain experts from six organizations (50 hours total) and six-phase thematic coding analysis. We propose the first annotation defect taxonomy for AI perception systems, comprising 18 error types organized along three dimensions: completeness, accuracy, and consistency. Analogous to Failure Mode and Effects Analysis (FMEA), this taxonomy serves as a “failure mode library” for annotation defects. Validated by industry practitioners, it supports root-cause analysis, supplier evaluation, annotator onboarding, and annotation guideline refinement—thereby enhancing AIePS development reliability and safety.
NLP data quality assessment has long relied on inter-annotator agreement, overlooking intra-annotator consistency—the temporal stability of individual annotators’ judgments. This neglect challenges the implicit “gold label as ground truth” assumption. Method: We conduct exploratory repeated annotation experiments across major NLP datasets and quantify intra-annotator agreement using Cohen’s and Fleiss’ Kappa, complemented by qualitative perceptual analysis. Contribution/Results: We demonstrate that mainstream NLP datasets routinely omit intra-annotator consistency reporting; moreover, individual annotators exhibit significant temporal variability in labeling identical texts. We identify and disentangle the dual influence of textual ambiguity and subjectivity on annotation stability. Our work establishes intra-annotator agreement as a foundational data quality metric, providing both a methodological framework and concrete guidelines for constructing more robust, reproducible NLP datasets.
This work addresses the lack of traceability in existing domain-specific fine-tuning approaches, which often leads to blind and inefficient data augmentation. The authors propose a “programming with data” paradigm that treats structured knowledge representations as a unified foundation for both training and evaluation, drawing an analogy to software development: training data serve as source code, model training as compilation, evaluation as unit testing, and data refinement as debugging. This framework enables precise, concept- and reasoning-chain–oriented model repair through structured knowledge extraction, test-driven data engineering, concept-level gap analysis, and diagnosis of broken reasoning chains. Validated across 16 disciplines, the approach significantly enhances model performance without compromising general capabilities, and the authors release an open-source knowledge base, evaluation suite, and training corpora to support reproducibility and further research.
This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.
This study addresses the inconsistency in human annotation caused by ambiguous category definitions in traditional content moderation. To resolve this, the authors propose an AI-driven constitutional annotation framework: large language models first assist humans in formulating structured, interpretable category “constitutions,” which then guide automated dual-axis labeling of intent and content safety. This approach shifts human effort from case-by-case judgments to high-level semantic definition. Evaluated on harassment, hate speech, and non-violent criminal conduct tasks, the method reduces cross-model annotation inconsistency by up to 57-fold compared to conventional paragraph-based rules and effectively exposes latent gaps in existing policy formulations.
This study addresses the lack of empirical evaluation regarding whether existing dataset documentation frameworks effectively foster developer reflectivity. Combining mixed-methods thematic analysis with corpus-assisted discourse analysis, the research systematically examines how prevailing documentation frameworks—and their real-world instantiations—cover core dimensions of reflectivity. The findings reveal, for the first time, that current frameworks consistently overlook critical reflective themes. Building on this insight, the authors develop a reflectivity-oriented coding manual and propose an enhanced datasheet template incorporating targeted prompts to elicit deeper reflection. This work offers actionable strategies and practical tools to strengthen the reflective capacity of dataset documentation practices.
Behavioral profiling (BP) annotation is challenging to automate due to its multidimensional, multilingual nature, and conventional task-level evaluation obscures underlying skill heterogeneity. This work proposes a novel “skill feasibility” paradigm, decomposing BP annotation into 14 operationalizable annotation skills and implementing a schema-guided, skill-document-driven pipeline. Evaluation over a 300-instance validation set—through two rounds of testing involving human annotators and large language models (GPT-5.4 and three open-source models)—reveals a “shared categorization, independent execution” pattern: humans and GPT exhibit high agreement at the skill level but diverge in instance-level execution. The study identifies five directly feasible skills, four recoverable via relabeling, and five structurally undefined. GPT-5.4 demonstrates reliable performance on feasible skills (accuracy = 0.678, κ = 0.665, weighted F1 = 0.695), whereas open-source models primarily fail in translating schemas into executable skills.