Score
Designs and implements annotation schemas, guidelines, and tooling and produces labeled text corpora by marking tokens, spans, relations, or corrections in textual data. Performs manual annotation and proofreading to ensure label consistency, resolve ambiguities, and document edge cases and annotation decisions.
High annotation costs and prolonged turnaround times plague NLP development, necessitating efficient and reliable data labeling paradigms. This paper proposes an LLM-powered Human-in-the-Loop (HITL) hybrid annotation framework that systematically integrates synthetic data generation, active learning, and human-AI collaboration, augmented with built-in mechanisms for annotation quality assessment, annotator management, and cost-benefit analysis. Unlike prior work—largely theoretical or narrowly scoped—this study introduces the first deployable, plug-and-play industrial-grade annotation methodology, bridging the critical gap between methodological research and real-world engineering practice. Empirical validation across multiple production NLP projects demonstrates that the framework consistently reduces annotation costs and cycle time by 30–50%, while maintaining label quality within required thresholds.
This study addresses the lack of systematic approaches for constructing, storing, and sharing high-quality annotated corpora. It proposes a generalizable and reusable end-to-end methodology encompassing annotation guideline development, corpus annotation, data storage, sharing mechanisms, and value realization, with an emphasis on full lifecycle management and cross-domain applicability. Integrating linguistic annotation theory, data management standards, and collaborative research practices, the approach is articulated through a structured framework and illustrative examples to yield a clear and actionable guide. The resulting methodology provides standardized support for diverse research domains, significantly enhancing the efficiency and quality with which researchers can build and utilize annotated textual data.
This work addresses the challenge that human-authored annotation guidelines are poorly suited for large language model (LLM)-based text annotation due to their informal, ambiguous, and context-dependent nature. We propose a guideline refactoring method oriented toward LLM auditing: automatically transforming natural-language guidelines into structured, semantically precise, instruction-style rules aligned with LLM comprehension preferences. Our approach preserves original semantic intent while systematically enhancing executability and robustness. Evaluated on disease entity recognition using the NCBI Disease Corpus, the refactored guidelines significantly improve LLM annotation accuracy and inter-annotator consistency, enabling automated iterative guideline refinement. Empirical analysis further identifies critical failure modes—including instruction ambiguity and insufficient coverage of edge cases. This study establishes a novel paradigm and reusable methodological framework for building high-quality, LLM-native annotation infrastructure.
NLP data quality assessment has long relied on inter-annotator agreement, overlooking intra-annotator consistency—the temporal stability of individual annotators’ judgments. This neglect challenges the implicit “gold label as ground truth” assumption. Method: We conduct exploratory repeated annotation experiments across major NLP datasets and quantify intra-annotator agreement using Cohen’s and Fleiss’ Kappa, complemented by qualitative perceptual analysis. Contribution/Results: We demonstrate that mainstream NLP datasets routinely omit intra-annotator consistency reporting; moreover, individual annotators exhibit significant temporal variability in labeling identical texts. We identify and disentangle the dual influence of textual ambiguity and subjectivity on annotation stability. Our work establishes intra-annotator agreement as a foundational data quality metric, providing both a methodological framework and concrete guidelines for constructing more robust, reproducible NLP datasets.
The exponential growth of academic literature poses significant challenges for efficiently constructing comparative tables in survey papers. Existing schema generation methods suffer from ambiguous evaluation criteria and limited editability. To address these issues, this paper proposes an intent-aware schema generation and editing framework: (1) it introduces intent modeling to mitigate semantic ambiguity in comparative dimension identification; (2) it designs an editable generation pipeline enabling on-demand customization of comparison dimensions; (3) it constructs the first benchmark dataset tailored for conditional schema generation; and (4) it integrates LLM-based prompt engineering with lightweight fine-tuning, combining one-shot generation and multi-stage editing strategies. Experimental results demonstrate that intent enhancement substantially improves schema reconstruction accuracy, while the editing mechanism further refines output quality. Notably, our lightweight fine-tuned model achieves performance competitive with state-of-the-art prompting-based large language models.
Educational annotation tools commonly lack scientifically grounded classification mechanisms, hindering effective organization of student annotations and their pedagogical utilization. Method: We systematically reviewed 32 educational annotation tools and, for the first time, proposed a four-category typology of classification paradigms: unclassified, controlled vocabulary, folksonomy, and ontology-driven classification. Building on literature analysis and conceptual framework development, we further introduced a four-dimensional classification model—comprehensive, interpretable, pedagogically aligned, and extensible. Contribution/Results: This model uncovers structural gaps in existing tools’ pedagogical support capabilities and establishes the first systematic classification benchmark and methodological foundation for the theoretical design and practical optimization of intelligent educational annotation systems.
This study addresses the lack of systematic, quantitative analysis regarding how document layout markers—such as paragraph boundary delimiters—in pretraining corpora influence language model behavior. The authors propose a novel “clean-window survival” metric and employ preregistered experiments, multi-model prediction difficulty assessments, and a reversible sidecar data format to empirically evaluate 13 public corpora. Their findings demonstrate that the critical factor affecting model predictions is not the specific delimiter symbols themselves, but rather the structural announcement information they convey. Removing these structural announcements substantially degrades model performance, whereas substituting the delimiter symbols has negligible impact. Moreover, models cannot spontaneously reconstruct missing structural cues. This work also provides the first quantitative characterization of layout marker distributions and introduces a “pure-frame” data format to disentangle markup from textual content.
This study addresses the widespread problem of incomplete reporting of annotation practices in natural language processing (NLP) research, which undermines reproducibility and quality assessment. Analyzing 1,603 papers from major NLP conferences between 2018 and 2025, the work introduces a unified taxonomy for annotation reporting that spans tasks, time, and domains, along with a minimal reporting standard. Leveraging a gold-standard dataset—Annotated-gold—curated through a combination of large language models and human adjudication, the authors construct Annotated-llm, achieving human-level inter-annotator agreement (Krippendorff’s α = 0.606) on structured information extraction. Despite gradual improvements in reporting over time, critical details—such as annotator training, linguistic competence, and compensation—remain frequently omitted. These findings advance the push toward more transparent and reliable annotation practices in NLP.
This study addresses the high cost and reliability challenges of employing large language models (LLMs) for structured annotation by proposing a dual-annotation-stream framework based on character-level alignment. The method automatically resolves unambiguous cases while routing conflicts to a browser-based interface for human adjudication. By integrating offline auditability, explicit logging strategies, and document- and span-level consistency computation, the framework supports field-level hybrid construction and direct export in original formats. Evaluated on a humanitarian benchmark, the system autonomously merges 8% of documents and precisely identifies 3,131 conflicts, substantially enhancing both review efficiency and result trustworthiness in human–machine collaborative annotation workflows.
This study addresses the challenging task of automatically identifying and segmenting legal conditions (Tatbestand) from legal consequences (Rechtsfolge) in German statutory texts. To facilitate research on this structural parsing problem, the authors introduce ANNOTARES, the first fine-grained annotated dataset covering three major German legal codes, enabling cross-code generalization studies. The work systematically evaluates a range of approaches, including rule-based baselines, CRF, BiLSTM, BiLSTM-CRF, and Transformer architectures based on BERT and large language models. Experimental results demonstrate that BERT-based and large language models significantly outperform traditional methods in capturing the complex syntactic structures inherent in legal texts, thereby confirming the effectiveness of pretrained language models for extracting logical structures in legal documents.
Existing approaches to automatic document formatting suffer from imprecise target localization and redundant content re-reading in content-aware scenarios, compounded by the absence of a dedicated evaluation benchmark. To address these limitations, this work introduces DocFormBench—the first comprehensive evaluation benchmark specifically designed for content-aware document formatting—and proposes DocFormFlow, a decoupled workflow that separates the task into two distinct phases: “what to format” (target localization) and “how to format” (format execution). By integrating large language models with multimodal models, DocFormFlow demonstrates significant improvements in formatting accuracy and substantially reduces token consumption across multiple mainstream models, underscoring precise target localization as a critical factor for high performance.