Gaze Behavior in Visual World Experiments Can be Modeled With Off-the-shelf Language-Vision Encoders

📅 2026-08-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the modeling of human predictive gaze behavior in visual world experiments, where participants integrate visual and linguistic inputs. The authors propose a method that leverages off-the-shelf CLIP-like multimodal dual encoders combined with bimodal attribution techniques to directly predict eye movement trajectories in a zero-shot setting—without any task-specific fine-tuning or generative architecture. This approach demonstrates for the first time that general-purpose multimodal encoders are sufficient to accurately reproduce the characteristic patterns of human predictive fixations observed in classic English visual world paradigms. The results provide strong evidence that such models inherently capture the alignment between language and vision necessary to predict real human gaze behavior, even without explicit training on eye-tracking data.
📝 Abstract
The recent advances in neural language models have also spurred much work in computational psycholinguistics, asking whether neural LMs are also promising models of human language processing. However, work has been overwhelmingly focused on the unimodal case of written or spoken language. In contrast, multimodal experimental paradigms, like visual world studies that present participants with both visual and linguistic input simultaneously, have been neglected. In this paper, we present a novel approach that predicts gaze behavior in visual world studies. It does so by combining a simple multi-modal bi-encoder model of the CLIP family with a bimodal attribution method. We demonstrate the ability of this approach to robustly replicate the results of a seminal English visual world study which shows hu- man predictive processing. Remarkably, it does so without a generative architecture and without the need for fine-tuning, despite not being trained for this task.
Problem

Research questions and friction points this paper is trying to address.

gaze behavior
visual world paradigm
multimodal modeling
computational psycholinguistics
language-vision encoders
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal modeling
gaze prediction
CLIP
visual world paradigm
computational psycholinguistics