Using LMs to Model the Effects of Context and Coreference during Sentence Comprehension

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the oversight in existing research regarding the role of long-distance discourse structure representations when strictly constraining language model context windows to simulate human working memory. Leveraging GPT-2 and large-scale naturalistic reading time data, this work systematically manipulates context length and employs counterfactual interventions to disrupt cross-sentence coreference chains, thereby investigating how long-range coreference tracking mechanisms influence cognitive alignment. The findings reveal a U-shaped relationship between context window size and psycholinguistic fit, with an optimal range of 500–1000 tokens. Furthermore, disrupting coreference chains degrades the predictive power of larger windows by 20%–40%. This research demonstrates that long-range coreference resolution constitutes a critical mechanism for enhancing the alignment between computational language models and human reading behavior.
📝 Abstract
Language models (LMs) are often used as a tool to model human language processing. Recent studies suggest that severely restricting LMs'context window improves their fit to human psycholinguistic data by simulating human working memory constraints. However, it is possible that this strict memory-decay approach overlooks humans'reliance on long-range structural representations, such as discourse structre. In this work, we systematically vary the context window size of GPT-2 across four large-scale naturalistic English reading-time datasets and observe a U-shaped relationship: Although restricted contexts (<20 tokens) successfully capture local memory limitations, expanded contexts (500--1,000 tokens) ultimately yield the highest overall psycholinguistic fit. To investigate the mechanism driving this benefit, we conduct a counterfactual inference-time experiment that disrupts cross-sentential entity chains by pronominalizing repeated discourse entities. Obscuring these structural linkages significantly degrades the predictive power of larger context windows by 20% to 40%. Our experiments demonstrate that tracking long-range coreference relations is one important factor for the alignment between LM surprisal and human reading behavior, and approximate the extent to which human comprehenders use global discourse relations during language processing.
Problem

Research questions and friction points this paper is trying to address.

language models
sentence comprehension
context window
coreference
discourse structure
Innovation

Methods, ideas, or system contributions that make the work stand out.

context window
coreference resolution
psycholinguistic modeling
counterfactual inference
discourse structure
🔎 Similar Papers
No similar papers found.