🤖 AI Summary
This study investigates the capacity of Transformer language models (GPT-2, LLaMA-7B, LLaMA2-7B) to model eye fixation durations during natural Spanish reading, evaluating their cognitive plausibility in capturing human linguistic predictability mechanisms. Method: Using eye-tracking data from Rioplatense Spanish readers, we employ word-level predictive probabilities from each model in regression analyses—constituting the first systematic assessment of open-source large language models’ cognitive interpretability in oculomotor prediction. Results: Transformers significantly outperform traditional n-gram and LSTM baselines, accounting for substantially more variance in fixation durations. However, even the best-performing models fail to capture the full variance explained by human predictability estimates, revealing a fundamental cognitive gap between current LMs and human neural prediction mechanisms. This work establishes a novel empirical benchmark bridging computational linguistics and cognitive neuroscience, offering theoretical insights into the limits of artificial systems as cognitive models of human language processing.
📝 Abstract
Recent advances in Natural Language Processing (NLP) have led to the development of highly sophisticated language models for text generation. In parallel, neuroscience has increasingly employed these models to explore cognitive processes involved in language comprehension. Previous research has shown that models such as N-grams and LSTM networks can partially account for predictability effects in explaining eye movement behaviors, specifically Gaze Duration, during reading. In this study, we extend these findings by evaluating transformer-based models (GPT2, LLaMA-7B, and LLaMA2-7B) to further investigate this relationship. Our results indicate that these architectures outperform earlier models in explaining the variance in Gaze Durations recorded from Rioplantense Spanish readers. However, similar to previous studies, these models still fail to account for the entirety of the variance captured by human predictability. These findings suggest that, despite their advancements, state-of-the-art language models continue to predict language in ways that differ from human readers.