STRUCTURALCOST: A controlled reading time dataset for modeling human sentence processing difficulty
This study addresses the scarcity of high-quality data isolating the processing costs of long-distance subject-verb dependencies, which impedes evaluating the cognitive plausibility of language models. To this end, we construct a self-paced reading dataset comprising 475 participants and 40,000 observations to quantify dependency resolution costs induced by syntactic embedding, systematically comparing the predictive performance of n-gram models, state space models (SSMs), and Transformers. By replicating low-powered psycholinguistic findings at NLP scale, we confirm that human reading times at the verb increase with dependency distance. Furthermore, our analysis reveals that current models only partially capture this difficulty gradient and consistently underestimate working memory integration costs, thereby establishing a critical benchmark for improving the cognitive alignment of computational language models.