Score
Designs and builds annotated datasets and corpora by defining annotation schemas and protocols, collecting and curating data, recruiting and instructing annotators (including native-speaker, bilingual, or multilingual annotators), and running annotation pipelines with quality-control to produce gold‑standard labeled resources. Produces and documents annotated resources across modalities (e.g., text corpora, image pixel masks, multi-/bi‑lingual datasets), releases datasets and benchmarks, and maintains annotation procedures and metadata for training and evaluation.
This study addresses the lack of systematic approaches for constructing, storing, and sharing high-quality annotated corpora. It proposes a generalizable and reusable end-to-end methodology encompassing annotation guideline development, corpus annotation, data storage, sharing mechanisms, and value realization, with an emphasis on full lifecycle management and cross-domain applicability. Integrating linguistic annotation theory, data management standards, and collaborative research practices, the approach is articulated through a structured framework and illustrative examples to yield a clear and actionable guide. The resulting methodology provides standardized support for diverse research domains, significantly enhancing the efficiency and quality with which researchers can build and utilize annotated textual data.
This work addresses the scarcity of high-quality coreference-annotated data and the prohibitive cost of manual annotation for coreference resolution. To this end, we propose a fully automated, dual-path data generation framework that requires no human labeling: (1) rule-based direct conversion leveraging existing corpora, and (2) cross-lingual joint parsing of dependency and coreference structures using multilingual pretrained models. Our framework is the first to systematically integrate structured data mapping with multilingual transfer capabilities, enabling coreference annotation for low-resource and unseen languages. Experiments across diverse languages demonstrate that the generated data achieves high quality and strong generalization, significantly reducing annotation costs. The approach establishes a scalable, reusable data infrastructure for coreference resolution—advancing both data efficiency and cross-lingual applicability in the field.
Low- and medium-resource language NLP faces critical challenges including data scarcity, insufficient cultural adaptation, and ethical violations in data annotation—such as platform exploitation of linguistic communities and disregard for annotator rights and welfare. Method: This study pioneers an integrated framework centered on “cultural embedding” and “labor dignity,” simultaneously foregrounding perspectives of language users and data workers. We employed a mixed-methods approach: multilingual community surveys (N=217), in-depth interviews (N=43), and critical discourse analysis coupled with thematic coding. Contribution/Results: The study yields 12 actionable, empirically grounded annotation guidelines. These have been adopted by three low-resource language NLP projects, resulting in a 47% increase in annotator retention and a 31% improvement in cultural adaptation scores (per expert evaluation). The framework advances both theoretical foundations and practical pathways for developing high-quality language resources that are linguistically authentic, culturally sensitive, and ethically sustainable.
Low-resource language benchmarks—exemplified by Turkish—frequently suffer from linguistic inaccuracy and cultural misalignment, undermining NLP evaluation validity. Method: We propose the first six-dimensional dataset quality framework for low-resource languages, assessing linguistic correctness, cultural appropriateness, terminological accuracy, among others, via expert human annotation and multi-model collaborative evaluation (GPT-4o, Llama3.3-70B). Contribution/Results: Systematic evaluation of 17 mainstream Turkish benchmarks reveals that 70% fail to meet baseline quality thresholds, and 85% of assessed dimensions exhibit significant deficiencies. LLMs substantially underperform humans on cultural commonsense reasoning, yet demonstrate complementary strengths: GPT-4o excels at syntactic and terminological judgment, while Llama3.3-70B outperforms on cultural knowledge inference. This work establishes a reproducible methodology and empirically grounded benchmark for rigorous low-resource language NLP evaluation.
Addressing the challenge of constructing high-quality, domain-specific annotated data—often costly and labor-intensive—this paper proposes a few-shot-driven synthetic data generation paradigm. Given only a small set of user-provided examples, the method retrieves semantically relevant real-world text from large-scale web corpora and leverages instruction-tuned large language models (LLMs) to automatically generate well-formatted, task-specific synthetic training data. It is the first approach to synergistically integrate corpus retrieval with LLM-based augmentation, enabling zero human annotation, domain adaptability, and efficient few-shot generalization. Empirical evaluation across biomedical, medical, and commonsense question answering (QA), as well as summarization tasks, demonstrates that models trained on the generated data achieve a 46-point preference score improvement over human-annotated baselines in summarization, while QA models match or surpass the performance of general-purpose foundation models.
This study addresses the widespread problem of incomplete reporting of annotation practices in natural language processing (NLP) research, which undermines reproducibility and quality assessment. Analyzing 1,603 papers from major NLP conferences between 2018 and 2025, the work introduces a unified taxonomy for annotation reporting that spans tasks, time, and domains, along with a minimal reporting standard. Leveraging a gold-standard dataset—Annotated-gold—curated through a combination of large language models and human adjudication, the authors construct Annotated-llm, achieving human-level inter-annotator agreement (Krippendorff’s α = 0.606) on structured information extraction. Despite gradual improvements in reporting over time, critical details—such as annotator training, linguistic competence, and compensation—remain frequently omitted. These findings advance the push toward more transparent and reliable annotation practices in NLP.
This work addresses the challenges of information retrieval (IR) for low-resource languages, where high-quality annotated data is scarce and automatically generated labels often suffer from reliability issues and biases. The authors propose BETA-Labeling, a novel framework that systematically evaluates the effectiveness of large language model (LLM)-assisted annotation in low-resource IR. By integrating multi-model collaborative labeling, context alignment, consistency verification, and majority voting—augmented with human evaluation—they construct the first high-quality Bengali IR dataset. The study also investigates the feasibility of reusing single-hop machine-translated data from other low-resource languages, revealing performance risks in cross-lingual transfer due to inconsistent semantic preservation and language-dependent biases. Experimental results demonstrate that the proposed approach substantially improves annotation quality, while the efficacy of cross-lingual data reuse is shown to be highly dependent on the linguistic characteristics of the language pair involved.
This study addresses the challenge of low-quality metadata that hinders dataset discoverability and reuse, particularly in the context of large language model (LLM)-generated descriptions lacking empirical guidance on context selection and its impact on quality. Building a literature-based framework for description quality assessment, the authors conduct systematic ablation experiments across 252 real-world CSV datasets. They uncover a previously unreported “table-structure penalty” phenomenon: relying solely on table structure significantly degrades narrative quality. While representative data samples aid semantic grounding, they do not improve overall human-rated quality. The work further reveals that different LLMs exhibit consistent descriptive styles. Through LLM-as-a-judge evaluations, semantic attribute analysis, and large-scale experimentation, the study offers key recommendations for LLM-assisted data publishing: concise, relevant context yields better results than redundant input, and table structure should be used cautiously as a basis for generation.
This study addresses the performance degradation commonly observed in multilingual large language models, which stems from imbalanced data distributions and the so-called “curse of multilinguality.” The authors identify the root cause as remediable corpus quality issues and propose a language-specific data curation and balancing strategy. By integrating multilingual quality evaluation with an efficient training mixture methodology, they optimize the composition of a 20-trillion-token corpus. Models trained on this refined dataset—specifically 3B and 8B parameter variants—achieve state-of-the-art multilingual performance while using 4–10 times fewer FLOPs than competing approaches. Furthermore, the curated corpus significantly enhances the multilingual scaling efficiency of Trinity Large (400B), demonstrating its effectiveness in improving both model performance and training efficiency across diverse languages.
This study addresses key challenges faced by social science researchers when using large language models (LLMs) for text annotation—namely, poor reproducibility, annotation errors that compromise statistical inference, and high technical barriers. To overcome these issues, the authors propose the first end-to-end LLM-based text annotation framework tailored specifically for the social sciences and humanities (SSH). The framework integrates structured prompt engineering, open-source LLM API integration, cross-validation, and error propagation modeling, with an explicit emphasis on avoiding prompt overfitting and quantifying annotation uncertainty. Implemented in both Python and R, this approach establishes a transparent, reproducible, and scalable workflow that substantially enhances the reliability, efficiency, and methodological rigor of automated text annotation in SSH research.