Similar Choices, Different Attention: Cross-Modal Associations in Humans and Vision-Language Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether the cross-modal alignment of Vision-Language Models (VLMs) extends beyond decision outcomes to replicate human gaze processes. For the first time, we synchronously collect human and VLM choice and eye-tracking data under identical stimuli, conducting a joint evaluation through multimodal experimental paradigms, attention mechanism analysis, and fine-tuning strategies. Our findings reveal that although VLMs occasionally match human choices, significant deviations persist in their fixation patterns. Furthermore, while fine-tuning improves decision accuracy, it fails to enhance gaze consistency. This work demonstrates that "correct outcomes" do not entail "similar cognitive processes," establishing that single-metric evaluations are insufficient for assessing cognitive alignment. Ultimately, this research provides a novel paradigm for evaluating VLMs by emphasizing the critical distinction between behavioral results and underlying attentional mechanisms.
📝 Abstract
Cross-modal associations are systematic pairings of features across modalities, such as the association of 'bouba' with round shapes and 'kiki' with sharp shapes. Prior work has compared humans and vision-language models (VLMs) on such associations, but often using different stimuli or tasks between humans and models. Here, we ask whether VLMs align with humans not only in choices, but also in where they look when making those choices. We study both VLMs and humans (N = 53), presenting them with the same stimuli, a pseudo-word and two images, and record participants' choices and eye movements, which we release. We find choice alignment in a few larger VLMs, but their saliency matches human gaze less closely than a center-bias baseline, a fixed Gaussian at the center of each image. Fine-tuning small VLMs on human choices brings their choice alignment to the level of a human majority-vote reference on unseen words and images, yet their attention still matches human gaze less closely than this baseline. Training model attention on human gaze raises attention-gaze correlation without improving choice alignment, and a single average gaze map per image position raises it by a similar amount. Matching human choices, or even human gaze patterns, is therefore not sufficient evidence of human-aligned cross-modal processing.
Problem

Research questions and friction points this paper is trying to address.

cross-modal associations
vision-language models
human alignment
eye movements
attention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-modal associations
Vision-language models
Eye tracking
Attention alignment
Center-bias baseline