surprisal analysis

Design and implement computations that assign per-token predictive surprisal (the negative log probability or information content) from sequence or language models, and produce analyses and visualizations of surprisal values across tokens. Use these measures to quantify model surprise at critical elements, correlate surprisal with human judgments, and diagnose or compare model predictions.

surprisalanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.19
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses a persistent source of predictive bias in surprisal theory: the ambiguity in defining linguistic units and their misalignment with tokenization schemes used by language models. To resolve this, the authors propose a unified framework that systematically disentangles unit definition from the selection of prediction regions—the first such formal separation in surprisal analysis. By treating tokenization as an implementation detail rather than a theoretical primitive, the approach integrates surprisal theory, the probabilistic mechanisms of pretrained language models, and formal modeling of linguistic units. This integration establishes a principled alignment between psycholinguistic experiments and computational models, substantially enhancing surprisal’s predictive validity, theoretical rigor, and cross-model comparability.

language modelslinguistic unitspredictability

This study critically examines the common practice of directly employing surprisal values from large language models (LLMs) to test the Surprisal Theory in psycholinguistics, highlighting that such usage overlooks the dependence of surprisal on model-specific representational and algorithmic choices. The work demonstrates that LLM-derived surprisal is not theoretically neutral but is significantly shaped by architectural differences—such as Transformer versus RNN—and algorithmic factors, including sampling strategies and context handling. Through three systematic comparative experiments, the authors reveal substantial variations in surprisal estimates across modeling decisions, thereby challenging its validity as a universal cognitive metric. The findings underscore the necessity for computational psycholinguistics to explicitly articulate representational and algorithmic commitments, advocating for more rigorous paradigms in theory validation.

computational-level theorylarge language modelspsycholinguistics

This work addresses the limitations of the traditional minimal pair paradigm, which is constrained by binary grammaticality judgments and reliance on text generation, making it ill-suited for characterizing model uncertainty. The authors propose a generation-free evaluation framework that extends surprisal-based (negative log-probability) analysis from binary contrasts to multi-domain ordinal classification tasks. By constructing full surprisal curves, the method simultaneously captures both model preferences and uncertainty, and—novelty introduced here—incorporates entropy to distinguish between ambiguous and unambiguous samples. Evaluated across four application domains, the approach consistently yields interpretable signals: surprisal curves exhibit pronounced minima at the expected rating levels, thereby validating the method’s effectiveness and generalizability.

language model evaluationminimal pairsmodel uncertainty

Testing the Predictions of Surprisal Theory in 11 Languages

Jul 07, 2023
EG
Ethan Gotlieb Wilcox
🏛️ ETH Zürich | University of Cambridge | MIT

Surprisal Theory has long been validated primarily in English native reading, limiting claims of its cross-linguistic universality. Method: This study systematically tests the theory across 11 typologically diverse languages spanning five major language families. Using monolingual and multilingual pretrained language models, we compute word-level surprisal and contextual entropy, then conduct hierarchical regression analyses against multilingual eye-tracking and reading-time data. Contribution/Results: We find robust linear relationships between surprisal and reading time, with contextual entropy exhibiting significant independent predictive power. Crucially, all three core theoretical predictions—(i) surprisal predicts reading time, (ii) entropy is predictive, and (iii) the surprisal–reading-time relationship is linear—are highly significant across all 11 languages. This establishes the broadest and most robust cross-linguistic link to date between information-theoretic measures and incremental language comprehension, providing the strongest multilingual empirical support for Surprisal Theory.

Examines surprisal-reading time relationship crosslinguisticallyTests Surprisal Theory in 11 diverse languagesValidates three key predictions of Surprisal Theory

On the Role of Context in Reading Time Prediction

Sep 12, 2024
AO
Andreas Opedal
🏛️ ETH Zürich | University of Zürich

Prior work has overestimated the role of surprisal in predicting reading times, conflating its effect with that of word frequency—a pervasive confound. Method: We introduce a novel orthogonal projection technique that projects language-model surprisal onto the null space of word frequency, thereby statistically isolating contextual predictability from lexical frequency. Combining surprisal with pointwise mutual information (PMI) and applying rigorous orthogonalization, we then quantify the unique variance in reading times explained solely by context via regression. Contribution/Results: Orthogonalized surprisal is uncorrelated with word frequency (r = 0), and its variance explained in reading times drops substantially—by 40–60%—relative to unadjusted surprisal. This demonstrates that context’s independent contribution to reading time prediction is markedly smaller than previously assumed. Our approach establishes a more rigorous, reproducible quantitative benchmark for evaluating contextual effects in language comprehension, challenging the dominant information-theoretic accounts grounded solely in surprisal.

Compares surprisal and PMI as contextual predictors in language modelsInvestigates how context influences reading time predictionProposes a frequency-orthogonalized predictor to isolate contextual effects

Latest Papers

What's happening recently
View more

Traditional surprisal quantifies language processing cost using a single scalar, overlooking the dynamic evolution of comprehension states. This work proposes “trajectory extrapolation error” as a novel metric: by fitting the historical trajectory of hidden states in Transformer language models (e.g., GPT-2, Pythia) and measuring the error in extrapolating this trajectory, it captures human sensitivity to local semantic momentum. This metric is orthogonal to surprisal and significantly predicts self-paced reading times on the Natural Stories corpus independently of surprisal, with particularly strong performance on garden-path sentences. The effect strengthens with model scale and replicates robustly across architectures, revealing that language processing cost comprises two dissociable components—prediction error and trajectory momentum.

garden-path sentenceshidden stateslanguage processing cost

Surprisal from language models is commonly employed as a proxy for metaphorical novelty, yet it is highly confounded with word frequency, potentially leading to attribution bias. This study systematically investigates the relationship between surprisal and metaphorical novelty by leveraging eight model sizes of Pythia, 154 training checkpoints, and two distinct word frequency measures, evaluated through regression analyses. The findings reveal that word frequency consistently emerges as a stronger predictor of novelty than surprisal. Moreover, the association between surprisal and novelty peaks early in training and rapidly diminishes thereafter, challenging prevailing assumptions about optimal surprisal mechanisms in language models and suggesting that prior results may have erroneously attributed frequency effects to contextual predictability.

contextual predictabilityfrequency confoundlanguage model surprisal

This study addresses the challenges of tokenization alignment and disputed prior validity in using language model perplexity to predict Chinese reading times. To overcome these obstacles, we propose a novel Shortest Matching Sequence (SMS) scheme for tokenization alignment. We conduct systematic statistical analyses leveraging from-scratch-trained Chinese-Pythia models (14M–1.4B parameters) across three Chinese eye-tracking corpora, including GECO-CN. Our findings confirm that perplexity effectively predicts Chinese reading times, thereby correcting previous null conclusions. Furthermore, we reveal that this predictive power is constrained by corpus characteristics and identify an inverse scaling phenomenon. This work cautions against deriving model scaling conclusions from single-corpus evaluations and establishes a new paradigm for computational cognitive modeling.

Chinese reading timeseye-tracking corporaLM surprisal

This work addresses the challenge of interpreting the contribution of input tokens to outputs in large language model generation by proposing the first model-agnostic probabilistic attribution method. The approach models text generation as a stochastic process and leverages Bayes’ rule to infer the conditional probability of a response given a prompt. Attribution scores are defined via the logarithm of probability ratios, while conditional entropy is introduced to quantify context sensitivity and generation uncertainty. Experiments across eight mainstream models and seven prompt categories demonstrate that the method effectively identifies anomalous generations, token-sensitive regions, and unstable behaviors, substantially enhancing users’ awareness and understanding of generative uncertainty.

interpretabilitylarge language modelsprobabilistic attribution

This study investigates the systematic underestimation by language models of cognitive load induced by syntactic ambiguities—such as garden-path sentences—during human reading. By modulating the number of parallel parse trees maintained in word-synchronous beam search within a recurrent neural network grammar (RNNG), the authors simulate surprisal under varying parsing capacities and use these estimates to predict eye-tracking reading times. This approach offers the first computational test of the “parsing multiplicity discrepancy hypothesis.” Results show that reducing the number of parallel parses amplifies the model’s prediction of garden-path effects, yet the magnitude remains substantially weaker than empirically observed human reading times, suggesting that limitations in parsing multiplicity alone cannot account for the mismatch between model surprisal and human cognitive processing difficulty.

garden path sentenceshuman sentence processinglanguage models

Hot Scholars

MG

Mario Giulianelli

Associate Professor, UCL
Computational LinguisticsLanguage ModellingAI Evaluation
BD

Byung-Doh Oh

New York University
computational linguisticspsycholinguisticsnatural language processing
EG

Ethan Gotlieb Wilcox

Asst. Prof. of Computational Linguistics @Georgetown. Previous: Postdoc @ETH, PhD @ Harvard
LinguisticsCognitive SciencePsycholinguisticsMachine Learning