🤖 AI Summary
Existing narrative research suffers from small-scale, coarse-grained, and poorly generalizable character typologies, lacking a reliable benchmark to evaluate models’ capacity for understanding dynamic character development. Method: We propose a fine-grained character attribute attribution task and introduce CHATTER—the first large-scale, manually annotated dataset of character attributes in film scripts (660 films, 2,998 characters, 12,967 attributes, 88,124 character–attribute pairs)—built via structured script parsing, cross-character–attribute semantic alignment, and rigorous quality control. A high-quality subset, CHATTEREVAL, validated through multi-round human annotation and inter-annotator agreement assessment, is released as a new evaluation benchmark. Contribution/Results: Empirical evaluation reveals substantial limitations of state-of-the-art language models in reasoning about temporally evolving character attributes. CHATTER provides a scalable, reproducible, and narratively grounded evaluation infrastructure for narrative understanding and long-context modeling.
📝 Abstract
Computational narrative understanding studies the identification, description, and interaction of the elements of a narrative: characters, attributes, events, and relations. Narrative research has given considerable attention to defining and classifying character types. However, these character-type taxonomies do not generalize well because they are small, too simple, or specific to a domain. We require robust and reliable benchmarks to test whether narrative models truly understand the nuances of the character's development in the story. Our work addresses this by curating the CHATTER dataset that labels whether a character portrays some attribute for 88124 character-attribute pairs, encompassing 2998 characters, 12967 attributes and 660 movies. We validate a subset of CHATTER, called CHATTEREVAL, using human annotations to serve as a benchmark to evaluate the character attribution task in movie scripts. evaldataset{} also assesses narrative understanding and the long-context modeling capacity of language models.