On the scaling relationship between cloze probabilities and language model next-token prediction

📅 2026-02-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how the scale of language models influences their ability to predict human cloze behavior, including eye movements and reading times. By evaluating next-word prediction performance across models of varying sizes against human cloze responses and cognitive metrics, the research finds that although larger models still systematically underestimate the probability of actual human responses, they significantly improve semantic alignment with human predictions due to enhanced memory and semantic modeling capabilities. Consequently, these models rely less on low-level lexical co-occurrence statistics. The work demonstrates that increasing model scale positively enhances cognitive modeling capacity, clarifying both the advantages and limitations of current large language models in predicting human language comprehension.

Technology Category

Natural Language Processing: (Large) Language ModelsMachine Learning: Large Multimodal Models (LMMs)Cognitive Modeling & Cognitive Systems: Simulating Human Behavior

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Recent work has shown that larger language models have better predictive power for eye movement and reading time data. While even the best models under-allocate probability mass to human responses, larger models assign higher-quality estimates of next tokens and their likelihood of production in cloze data because they are less sensitive to lexical co-occurrence statistics while being better aligned semantically to human cloze responses. The results provide support for the claim that the greater memorization capacity of larger models helps them guess more semantically appropriate words, but makes them less sensitive to low-level information that is relevant for word recognition.
Problem

Research questions and friction points this paper is trying to address.

cloze probability
language model
next-token prediction
semantic alignment
lexical co-occurrence
Innovation

Methods, ideas, or system contributions that make the work stand out.

scaling relationship
cloze probability
language models
semantic alignment
next-token prediction
🔎 Similar Papers
No similar papers found.
C
Cassandra L. Jacobs
Department of Linguistics, University at Buffalo
M
Morgan Grobol
MoDyCo, Université Paris Nanterre