narrative annotation design

Designs and implements annotation schemes and coding protocols to label and extract interpretable dimensions of narrative text (e.g., agency, setting, events). This includes creating coding guidelines, sampling strategies for passages, and procedures for reliable labeling and quality control.

narrativeannotationdesign

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study investigates whether BERT embeddings implicitly encode multidimensional semantic structures—namely temporal, spatial, causal, and character-related information—in fictional narratives. The authors construct the first token-level narrative annotation dataset across these four dimensions with the aid of large language models and systematically evaluate BERT’s representational capacity using linear probing, class-balanced weighting, confusion matrices, and clustering analysis assessed via Adjusted Rand Index (ARI). Results demonstrate that BERT indeed captures significant narrative information, achieving 94% accuracy and a macro-averaged recall of 0.83 with linear probes. Causal (0.75) and spatial (0.66) dimensions exhibit particularly strong performance, significantly outperforming random baselines. The study also identifies a “boundary leakage” phenomenon, revealing that while narrative semantics are present, they do not form discrete clusters—providing the first token-level evidence of BERT’s implicit encoding of complex narrative structures.

BERT embeddingscausalitynarrative semantics

A Decomposition-Based Approach for Evaluating Inter-Annotator Disagreement in Narrative Analysis

Jun 11, 2022
EL
Effi Levi
🏛️ The Hebrew University of Jerusalem

This study addresses the poorly understood sources of inter-annotator disagreement in narrative analysis. We propose decomposing the annotation task into two hierarchical subtasks: (1) judgment of narrative event existence and (2) identification of narrative elements. This constitutes the first structured decomposition and quantitative attribution of annotation disagreement in narrative analysis. Using a sentence-level dataset annotated for Complication, Resolution, and Success, we apply conceptual decomposition, variance decomposition, and qualitative case analysis. Results reveal significantly divergent contributions of the two subtasks to overall disagreement. Further analysis identifies three actionable root causes: textual ambiguity, ill-defined schema criteria, and annotator subjectivity. The proposed framework provides a principled, operational diagnostic tool and concrete pathways for improving annotation consistency in narrative analysis.

Analyze inter-annotator disagreement by decomposing annotation levelsCompare theoretically-driven and exploration-based decomposition strategiesReveal latent structures and sources of annotation disagreements

This study addresses the lack of systematic approaches for constructing, storing, and sharing high-quality annotated corpora. It proposes a generalizable and reusable end-to-end methodology encompassing annotation guideline development, corpus annotation, data storage, sharing mechanisms, and value realization, with an emphasis on full lifecycle management and cross-domain applicability. Integrating linguistic annotation theory, data management standards, and collaborative research practices, the approach is articulated through a structured framework and illustrative examples to yield a clear and actionable guide. The resulting methodology provides standardized support for diverse research domains, significantly enhancing the efficiency and quality with which researchers can build and utilize annotated textual data.

annotated corpusannotation guidelinescorpus creation

This study addresses the lack of high-quality, structured annotation methods for narrative structures of economic events in news discourse, which hinders natural language processing models from effectively capturing complex causal relationships. To bridge this gap, the authors propose a directed acyclic graph (DAG)-based narrative annotation framework that integrates qualitative content analysis, where nodes represent events and directed edges encode causal relations. The framework incorporates local structural constraints—such as one-hop neighborhood restrictions—to reduce inter-annotator variability. Through a 6×3 factorial experiment, the authors systematically evaluate how different graph representations and distance metrics affect annotation reliability, finding that lenient metrics tend to overestimate agreement, whereas local constraints significantly improve Krippendorff’s α. The project publicly releases an implementation of α adapted for graph-structured data, offering a robust methodological foundation and practical guidance for narrative graph annotation.

graph-based representationhuman label variationinter-annotator agreement

Latest Papers

What's happening recently
View more

This study addresses the overreliance on inter-annotator agreement in current data annotation practices, which often overlooks annotation’s capacity to capture conceptual validity as a measurement act. Treating annotation as a measurement process, the work identifies five root causes of annotation issues—errors, ambiguity, impossibility, subjectivity, and annotator identity—and develops a measurement theory–based framework for diagnosing and improving annotation quality. Drawing on a synthesis of 132 literature sources and 10 semi-structured interviews, the research systematically defines target constructs, designs annotation instruments, implements labeling procedures, and evaluates both reliability and validity. The resulting framework equips annotation teams with evaluation methods that transcend mere agreement metrics, thereby substantially strengthening the foundational quality of AI training data.

annotation qualitydata annotationmeasurement

This study addresses the challenges of quality assessment and contamination detection in large model training data, where absolute ground truth is often unavailable. We introduce the concept of "internal annotator dynamics," which analyzes a single annotator's self-reproducibility consistency on identical texts over time. Using the segmentation and causal-functional annotation of Sumerian mythology as an experimental testbed, we compare annotations produced by human experts and large language models (LLMs) across different time points. Our findings reveal that human annotators exhibit a distinctive pattern characterized by "stable boundaries but drifting labels." We demonstrate that this temporal signature can serve as an effective metric for data quality assessment and contamination detection, enabling the reliable identification of annotations that appear human-generated but are actually produced by LLMs.

annotation consistencydata qualityintra-annotator dynamics

This study addresses the lack of systematic, quantitative analysis regarding how document layout markers—such as paragraph boundary delimiters—in pretraining corpora influence language model behavior. The authors propose a novel “clean-window survival” metric and employ preregistered experiments, multi-model prediction difficulty assessments, and a reversible sidecar data format to empirically evaluate 13 public corpora. Their findings demonstrate that the critical factor affecting model predictions is not the specific delimiter symbols themselves, but rather the structural announcement information they convey. Removing these structural announcements substantially degrades model performance, whereas substituting the delimiter symbols has negligible impact. Moreover, models cannot spontaneously reconstruct missing structural cues. This work also provides the first quantitative characterization of layout marker distributions and introduces a “pure-frame” data format to disentangle markup from textual content.

boundary inferencedataset documentationdocument structure

Hot Scholars

DK

Dongyeop Kang

University of Minnesota
Natural Language Processing
KD

Karin de Langis

PhD Candidate, University of Minnesota
Artificial IntelligenceRoboticsComputer Vision
IP

Ivan P. Yamshchikov

Research Professor at CAIRO, THWS
natural language generationcomputational creativityempathetic aiethics of ai application
ZW

Zhijing Wu

Beijing Institute of Technology
Information RetrievalNatural Language Processing
GR

Giuseppe Russo

Dipartimento di Ingegneria Chimica, Gestionale, Informatica e Meccanica
Knowledge Engineering