Score
Computing surprisal (negative log probability) and incremental probability updates under probabilistic models — including syntactic or lexical surprisal — to predict processing difficulty, acceptability, and reading measures.
Traditional lexical surprisal struggles to effectively predict processing difficulty in garden-path sentences. This work proposes a syntactic-structure-based measure of processing difficulty that does not rely on lexical probabilities. Specifically, it employs incremental syntactic parsing to dynamically maintain a probability distribution over syntactic trees—termed “syntactic belief”—and quantifies the magnitude of belief updating at critical ambiguity points using generalized Rényi divergence as a proxy for cognitive load. This approach significantly outperforms traditional lexical surprisal in fitting human reading time data, offering psycholinguistics a non-lexical, syntax-driven alternative metric grounded in belief updating. The findings open a new avenue for understanding online language processing through the lens of syntactic uncertainty resolution.
Surprisal Theory has long been validated primarily in English native reading, limiting claims of its cross-linguistic universality. Method: This study systematically tests the theory across 11 typologically diverse languages spanning five major language families. Using monolingual and multilingual pretrained language models, we compute word-level surprisal and contextual entropy, then conduct hierarchical regression analyses against multilingual eye-tracking and reading-time data. Contribution/Results: We find robust linear relationships between surprisal and reading time, with contextual entropy exhibiting significant independent predictive power. Crucially, all three core theoretical predictions—(i) surprisal predicts reading time, (ii) entropy is predictive, and (iii) the surprisal–reading-time relationship is linear—are highly significant across all 11 languages. This establishes the broadest and most robust cross-linguistic link to date between information-theoretic measures and incremental language comprehension, providing the strongest multilingual empirical support for Surprisal Theory.
This study critically examines the common practice of directly employing surprisal values from large language models (LLMs) to test the Surprisal Theory in psycholinguistics, highlighting that such usage overlooks the dependence of surprisal on model-specific representational and algorithmic choices. The work demonstrates that LLM-derived surprisal is not theoretically neutral but is significantly shaped by architectural differences—such as Transformer versus RNN—and algorithmic factors, including sampling strategies and context handling. Through three systematic comparative experiments, the authors reveal substantial variations in surprisal estimates across modeling decisions, thereby challenging its validity as a universal cognitive metric. The findings underscore the necessity for computational psycholinguistics to explicitly articulate representational and algorithmic commitments, advocating for more rigorous paradigms in theory validation.
This study addresses the limited ability of current neural language models to predict human reading difficulty in garden-path sentences using surprisal, which fails to account for the associated cognitive load. To bridge this gap, the authors propose fine-tuning off-the-shelf neural language models on human reading time data from garden-path constructions, using supervised optimization to align model surprisal more closely with empirical reading behavior. The experiments demonstrate, for the first time, that such fine-tuned models can simultaneously explain both garden-path effects and reading times in naturalistic contexts without compromising general language modeling performance. This approach significantly improves prediction accuracy across both types of data, offering a unified account of human sentence processing grounded in modern language models.
Prior work has overestimated the role of surprisal in predicting reading times, conflating its effect with that of word frequency—a pervasive confound. Method: We introduce a novel orthogonal projection technique that projects language-model surprisal onto the null space of word frequency, thereby statistically isolating contextual predictability from lexical frequency. Combining surprisal with pointwise mutual information (PMI) and applying rigorous orthogonalization, we then quantify the unique variance in reading times explained solely by context via regression. Contribution/Results: Orthogonalized surprisal is uncorrelated with word frequency (r = 0), and its variance explained in reading times drops substantially—by 40–60%—relative to unadjusted surprisal. This demonstrates that context’s independent contribution to reading time prediction is markedly smaller than previously assumed. Our approach establishes a more rigorous, reproducible quantitative benchmark for evaluating contextual effects in language comprehension, challenging the dominant information-theoretic accounts grounded solely in surprisal.
This study addresses a persistent source of predictive bias in surprisal theory: the ambiguity in defining linguistic units and their misalignment with tokenization schemes used by language models. To resolve this, the authors propose a unified framework that systematically disentangles unit definition from the selection of prediction regions—the first such formal separation in surprisal analysis. By treating tokenization as an implementation detail rather than a theoretical primitive, the approach integrates surprisal theory, the probabilistic mechanisms of pretrained language models, and formal modeling of linguistic units. This integration establishes a principled alignment between psycholinguistic experiments and computational models, substantially enhancing surprisal’s predictive validity, theoretical rigor, and cross-model comparability.
Traditional surprisal quantifies language processing cost using a single scalar, overlooking the dynamic evolution of comprehension states. This work proposes “trajectory extrapolation error” as a novel metric: by fitting the historical trajectory of hidden states in Transformer language models (e.g., GPT-2, Pythia) and measuring the error in extrapolating this trajectory, it captures human sensitivity to local semantic momentum. This metric is orthogonal to surprisal and significantly predicts self-paced reading times on the Natural Stories corpus independently of surprisal, with particularly strong performance on garden-path sentences. The effect strengthens with model scale and replicates robustly across architectures, revealing that language processing cost comprises two dissociable components—prediction error and trajectory momentum.
This study investigates the systematic underestimation by language models of cognitive load induced by syntactic ambiguities—such as garden-path sentences—during human reading. By modulating the number of parallel parse trees maintained in word-synchronous beam search within a recurrent neural network grammar (RNNG), the authors simulate surprisal under varying parsing capacities and use these estimates to predict eye-tracking reading times. This approach offers the first computational test of the “parsing multiplicity discrepancy hypothesis.” Results show that reducing the number of parallel parses amplifies the model’s prediction of garden-path effects, yet the magnitude remains substantially weaker than empirically observed human reading times, suggesting that limitations in parsing multiplicity alone cannot account for the mismatch between model surprisal and human cognitive processing difficulty.
Surprisal from language models is commonly employed as a proxy for metaphorical novelty, yet it is highly confounded with word frequency, potentially leading to attribution bias. This study systematically investigates the relationship between surprisal and metaphorical novelty by leveraging eight model sizes of Pythia, 154 training checkpoints, and two distinct word frequency measures, evaluated through regression analyses. The findings reveal that word frequency consistently emerges as a stronger predictor of novelty than surprisal. Moreover, the association between surprisal and novelty peaks early in training and rapidly diminishes thereafter, challenging prevailing assumptions about optimal surprisal mechanisms in language models and suggesting that prior results may have erroneously attributed frequency effects to contextual predictability.
This study examines the falsifiability of the relationship between human language processing difficulty and surprisal as predicted by language models. Through formal analysis, information-theoretic reasoning, and psycholinguistic modeling, it demonstrates that surprisal theory, absent additional constraints on the language model, collapses into a tautology and fails to yield testable predictions. The authors prove that any non-negative processing difficulty profile can be linearly approximated by the surprisal of some language model, revealing a fundamental flaw in the theory’s implicit empiricist assumptions. To overcome this tautological impasse, the paper advocates for integrating cognitive process models that are independent of behavioral data, thereby establishing a rationalist framework with genuine explanatory power.