🤖 AI Summary
This study addresses the scarcity of high-quality data isolating the processing costs of long-distance subject-verb dependencies, which impedes evaluating the cognitive plausibility of language models. To this end, we construct a self-paced reading dataset comprising 475 participants and 40,000 observations to quantify dependency resolution costs induced by syntactic embedding, systematically comparing the predictive performance of n-gram models, state space models (SSMs), and Transformers. By replicating low-powered psycholinguistic findings at NLP scale, we confirm that human reading times at the verb increase with dependency distance. Furthermore, our analysis reveals that current models only partially capture this difficulty gradient and consistently underestimate working memory integration costs, thereby establishing a critical benchmark for improving the cognitive alignment of computational language models.
📝 Abstract
We introduce STRUCTURALCOST, a self-paced reading dataset of 475 participants and 40,800 observations isolating the processing cost of long-distance subject-verb dependency resolution. We replicate a low-powered psycholinguistic finding at NLP scale, namely that human reading times at the main verb increase with dependency length, driven by syntactic embedding beyond linear distance. Different language models -- spanning n-gram models, SSMs, and transformers -- partially mirror this graded difficulty profile, yet underestimate the integration cost humans incur, with a gap that persists across architectures and model sizes. This suggests these models capture the predictive component of human processing but not the full integration cost that working memory imposes. STRUCTURALCOST provides data needed to drive progress toward evaluating the cognitive plausibility of language models.