Score
Designs and implements automated annotation and classification systems that assign standardized codes or labels to content items using rule-based, statistical, or LLM-based models, including development of codebooks, preprocessing and labeling pipelines, and human-in-the-loop workflows. Builds validation and evaluation processes to compare automated labels to expert coding and to analyze label outputs for prevalence, co-occurrence, and temporal patterns.
This work systematically evaluates the efficacy and limitations of large language models (LLMs) as annotators for subjective tasks. We survey 12 prior studies and empirically compare opinion distribution alignment between GPT-series models and human annotators across four subjective datasets—introducing, for the first time, a novel evaluation paradigm centered on *perspective diversity alignment*. Results reveal substantial distributional biases in LLMs, including underrepresentation of minority viewpoints, prompt sensitivity, English-language bias, and embedded societal prejudices; most existing annotation methods overlook such distributional discrepancies, and only a few strategies effectively capture opinion diversity. Our findings expose critical reliability risks in deploying LLMs for subjective annotation and establish a reproducible statistical framework—comprising quantitative metrics and methodological guidelines—for assessing annotation quality in subjective NLP tasks.
Existing fine-grained corpus analysis relies either on labor-intensive manual annotation or opaque statistical methods, compromising scalability and interpretability. This paper proposes a two-stage inductive coding framework powered by large language models (LLMs): (1) bottom-up prompt engineering for automated, fine-grained label generation; and (2) semantic embedding–guided hierarchical clustering to construct interpretable, multi-level topic structures. Grounded in qualitative research logic, the approach ensures analytical transparency and human controllability while substantially enhancing scalability for large-scale qualitative text analysis. Experiments across three heterogeneous datasets demonstrate high alignment with expert annotations (average F1 = 0.82). Applied to opioid litigation texts, the method successfully uncovered systemic, aggressive marketing strategies employed by pharmaceutical companies—validating its theoretical insightfulness and practical utility.
This study addresses key challenges faced by social science researchers when using large language models (LLMs) for text annotation—namely, poor reproducibility, annotation errors that compromise statistical inference, and high technical barriers. To overcome these issues, the authors propose the first end-to-end LLM-based text annotation framework tailored specifically for the social sciences and humanities (SSH). The framework integrates structured prompt engineering, open-source LLM API integration, cross-validation, and error propagation modeling, with an explicit emphasis on avoiding prompt overfitting and quantifying annotation uncertainty. Implemented in both Python and R, this approach establishes a transparent, reproducible, and scalable workflow that substantially enhances the reliability, efficiency, and methodological rigor of automated text annotation in SSH research.
This study addresses the high false positive rates of large language models (LLMs) in zero- or few-shot qualitative coding, particularly stemming from “definition misinterpretation” and “meta-discussion confusion,” which severely degrade performance on low-accuracy categories. To mitigate this, the authors propose a two-stage automated coding framework: in the first stage, an LLM applies labels based on a human-curated codebook; in the second, a critic LLM performs targeted self-reflection by re-evaluating positive labels using both the original text and the initial reasoning trace. Guided by human validation, the approach incorporates codebook adaptation, specialized critique rules, and F1-oriented evaluation. Evaluated on a dataset of 3,000 emails across six coding categories, the method improves F1 scores by 0.04–0.25, with two previously low-performing categories rising significantly from 0.52/0.55 to 0.69/0.79, effectively suppressing noise and restoring classifier usability.
High annotation costs and prolonged turnaround times plague NLP development, necessitating efficient and reliable data labeling paradigms. This paper proposes an LLM-powered Human-in-the-Loop (HITL) hybrid annotation framework that systematically integrates synthetic data generation, active learning, and human-AI collaboration, augmented with built-in mechanisms for annotation quality assessment, annotator management, and cost-benefit analysis. Unlike prior work—largely theoretical or narrowly scoped—this study introduces the first deployable, plug-and-play industrial-grade annotation methodology, bridging the critical gap between methodological research and real-world engineering practice. Empirical validation across multiple production NLP projects demonstrates that the framework consistently reduces annotation costs and cycle time by 30–50%, while maintaining label quality within required thresholds.
This study investigates whether large language models (LLMs) can reliably replace human annotators to mitigate the high cost and logistical complexity of human-subject studies in software engineering innovation evaluation. We systematically evaluate six state-of-the-art LLMs across ten code-related annotation tasks—including code summary quality assessment and defect repair judgment—using five public datasets. Methodologically, we propose *inter-model agreement* as a novel task-adaptability predictor and integrate confidence-threshold filtering to identify samples safe for LLM-only annotation, thereby establishing a hybrid human–LLM evaluation paradigm. Results show that LLMs achieve or approach human inter-annotator agreement (Krippendorff’s α ≥ 0.8) on multiple tasks; inter-model agreement strongly predicts task feasibility (AUC = 0.92); and confidence-based filtering raises replacement accuracy to 94.3%.
This study addresses the inconsistency in human annotation caused by ambiguous category definitions in traditional content moderation. To resolve this, the authors propose an AI-driven constitutional annotation framework: large language models first assist humans in formulating structured, interpretable category “constitutions,” which then guide automated dual-axis labeling of intent and content safety. This approach shifts human effort from case-by-case judgments to high-level semantic definition. Evaluated on harassment, hate speech, and non-violent criminal conduct tasks, the method reduces cross-model annotation inconsistency by up to 57-fold compared to conventional paragraph-based rules and effectively exposes latent gaps in existing policy formulations.
This study addresses the scalability challenges of traditional manual coding when analyzing large-scale textual data generated from K–12 teachers’ interactions with AI systems. It proposes a human-AI collaborative qualitative analysis framework that positions large language models (LLMs) as assistive annotation tools rather than interpretive agents, while preserving human researchers’ central role in defining concepts and constructing analytical frameworks through open, axial, and selective coding. Key innovations include an auditable human-AI workflow, a set-valued intercoder agreement metric suited for multi-label contexts, and a multi-round calibration mechanism. The resulting codebook comprises 72 codes organized into 19 categories across six domains, demonstrating high reliability on a dataset of 2,560 messages. The approach’s validity and scalability are further evidenced by human researchers identifying five novel codes missed by the LLM.
This work addresses the challenges large language models face in generating accurate and readable statistical visualizations, stemming from existing datasets' lack of full alignment among code, data context, and question-answer pairs. The authors propose a structured, multi-stage workflow that decomposes chart generation into verifiable steps—data filtering, plotting proposal, code synthesis, rendering, and validation-driven refinement—and introduces a rendering feedback mechanism to transform the task from one-shot code generation into an iterative, verifiable process. The approach jointly produces charts, code, contextual metadata, natural language descriptions, and associated question-answer pairs, yielding a benchmark of 1,500 visualizations (spanning 24 chart types) and 30,003 QA pairs across 74 UCI datasets. Evaluation of 16 vision-language models demonstrates that this framework effectively exposes their limitations in numerical extraction, comparison, and reasoning, enabling fine-grained assessment of visual reasoning capabilities.
This study presents the first systematic evaluation of large language models’ (LLMs’) capability to perform automated qualitative coding in cybersecurity contexts, aiming to replace costly expert human annotation. Using the LiveBench platform, four state-of-the-art LLMs were prompted with realistic strategies—including detailed coding manuals, exemplar guidance, and conflicting examples—to code free-text participant comments on vulnerable code according to security-relevant categories. Inter-rater agreement between model-generated and human annotations was assessed using Cohen’s Kappa. Results indicate that LLM performance improves only marginally when provided with detailed code descriptions, yet remains inconsistent overall, falling short of reliably substituting human coders. These findings highlight current limitations in LLMs’ ability to comprehend nuanced technical contexts inherent in security-related qualitative analysis.