🤖 AI Summary
This study addresses the modeling of human predictive gaze behavior in visual world experiments, where participants integrate visual and linguistic inputs. The authors propose a method that leverages off-the-shelf CLIP-like multimodal dual encoders combined with bimodal attribution techniques to directly predict eye movement trajectories in a zero-shot setting—without any task-specific fine-tuning or generative architecture. This approach demonstrates for the first time that general-purpose multimodal encoders are sufficient to accurately reproduce the characteristic patterns of human predictive fixations observed in classic English visual world paradigms. The results provide strong evidence that such models inherently capture the alignment between language and vision necessary to predict real human gaze behavior, even without explicit training on eye-tracking data.
📝 Abstract
The recent advances in neural language models have also spurred much work in computational psycholinguistics, asking whether neural LMs are also promising models of human language processing. However, work has been overwhelmingly focused on the unimodal case of written or spoken language. In contrast, multimodal experimental paradigms, like visual world studies that present participants with both visual and linguistic input simultaneously, have been neglected. In this paper, we present a novel approach that predicts gaze behavior in visual world studies. It does so by combining a simple multi-modal bi-encoder model of the CLIP family with a bimodal attribution method. We demonstrate the ability of this approach to robustly replicate the results of a seminal English visual world study which shows hu- man predictive processing. Remarkably, it does so without a generative architecture and without the need for fine-tuning, despite not being trained for this task.