Encoding EEG Signals to Examine Human-Like Next-Word Prediction Behaviour in Language Models

📅 2026-07-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether language models, while improving next-word prediction performance, genuinely emulate the cognitive-neural mechanisms underlying human reading. By constructing information-theoretic regressors based on top-1 prediction accuracy and surprisal, the authors pioneer the use of electroencephalographic event-related potentials (ERPs) for fine-grained evaluation of alignment between language models and human language processing. The findings reveal that only surprisal significantly predicts ERP components associated with linguistic processing, particularly for open-class words carrying high semantic load. Moreover, increasing model scale does not necessarily enhance alignment with human neural signals, thereby challenging the prevailing assumption that larger models are inherently more human-like in their cognitive fidelity.
📝 Abstract
Language models (LMs) are trained to excel at predicting the next word in the sequence given prior context, and humans also share this predictability in reading comprehension. Neuroscience research reveals that next-word predictability influences brain response, as recorded at millisecond resolution using electroencephalography (EEG). While our evidence indicates that advanced LMs achieve accuracies closely aligned with human performance at the next-word prediction task, this raises the question: Does higher prediction accuracy necessarily mean that these models adequately capture the cognitive signals associated with human reading comprehension? Here, we generate regressors for both humans and LMs based on two information measures, including top-1 prediction and surprisal, to predict event-related potential (ERP) elicited from EEG recordings which reflect different stages of cognitive processing during reading. We argue that modelling ERP patterns offers fine-grained analysis of the cognitive plausibility of various LMs during reading. Our results indicate that only surprisal potentially correlates with language-processing ERPs, especially for open-class words with high semantic content. Moreover, our findings challenge the assumption that scaling LMs with increased parameters and computational budgets will consistently lead to improved convergence with human-like linguistic processing.
Problem

Research questions and friction points this paper is trying to address.

next-word prediction
EEG
event-related potential
cognitive plausibility
language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

surprisal
event-related potential
EEG
language models
cognitive plausibility
🔎 Similar Papers
No similar papers found.