🤖 AI Summary
This study investigates how the scale of language models influences their ability to predict human cloze behavior, including eye movements and reading times. By evaluating next-word prediction performance across models of varying sizes against human cloze responses and cognitive metrics, the research finds that although larger models still systematically underestimate the probability of actual human responses, they significantly improve semantic alignment with human predictions due to enhanced memory and semantic modeling capabilities. Consequently, these models rely less on low-level lexical co-occurrence statistics. The work demonstrates that increasing model scale positively enhances cognitive modeling capacity, clarifying both the advantages and limitations of current large language models in predicting human language comprehension.
📝 Abstract
Recent work has shown that larger language models have better predictive power for eye movement and reading time data. While even the best models under-allocate probability mass to human responses, larger models assign higher-quality estimates of next tokens and their likelihood of production in cloze data because they are less sensitive to lexical co-occurrence statistics while being better aligned semantically to human cloze responses. The results provide support for the claim that the greater memorization capacity of larger models helps them guess more semantically appropriate words, but makes them less sensitive to low-level information that is relevant for word recognition.