Score
Designs, builds, and evaluates models that predict or generate human scanpaths—sequences of fixations and saccades, including reading scanpaths—conditioned on prompts, task instructions, or viewer identity. Implements prompt-based conditioning and related techniques to adapt predicted eye-movement behavior without changing model architecture.
Existing scanpath prediction methods predominantly rely on population-level models, neglecting inter-individual heterogeneity in eye movements and thus limiting applicability to real-world social human–computer interaction. Method: To address individual gaze prediction in dynamic social videos, we propose the first unified deep learning model that jointly encodes fixation history and social cues via a gated recurrent fusion mechanism, augmented with sequence-wise attention for dynamic saliency representation. The model implicitly learns both universal attention mechanisms and subject-specific patterns within a single architecture. Contribution/Results: Evaluated on a free-viewing social video dataset, it achieves performance on par with or superior to subject-specific models. Large-scale experiments demonstrate that late fusion significantly outperforms early fusion, validating the efficacy of combining universal representations with personalized supervision. Notably, this is the first work to enable cross-subject generalization—predicting diverse observers’ scanpaths using a single model—thereby substantially improving inter-subject transferability.
This work investigates whether readers’ reading goals—information seeking versus comprehension reading—can be decoded in real time from their eye movement trajectories. To address this, we introduce the first large-scale, annotated eye-tracking dataset and propose a novel Transformer architecture that jointly models scanpath representations and contextualized semantic features from pretrained language models; we further incorporate a mixed-effects model to disentangle text difficulty from inter-subject variability. Methodologically, our approach innovatively integrates sequential visual behavior with linguistic semantics and employs ensemble learning to enhance robustness. Experiments demonstrate that the model achieves high-accuracy goal classification early in reading—well before completion—and identify key determinants of classification difficulty, including textual features (e.g., information density, structural complexity) and individual differences. These findings substantially advance the understanding of oculomotor correlates underlying distinct cognitive mechanisms in the two reading modes.
This work addresses the limitations of traditional scanpath modeling, which relies on handcrafted architectures and struggles to flexibly incorporate conditional information such as task instructions or individual differences. For the first time, the authors introduce a vision-language model to this task by discretizing gaze coordinates into textual tokens and formulating scanpath prediction as an autoregressive sequence generation problem. Through prompt engineering, the framework jointly models multiple factors—including viewer identity, task goals, and fixation durations—within a unified architecture. The approach not only enables precise computation of information gain per fixation but also offers potential for behavioral intervention. Evaluated on MIT1003, it achieves an information gain of 2.18 bits, a 46% improvement over DeepGaze III, establishes new state-of-the-art results across multiple benchmarks, and successfully reproduces known eye movement phenomena.
To address the challenge of personalized scanpath prediction under few-shot settings, this paper introduces the novel Few-Shot Personalized Scanpath Prediction (FS-PSP) task. To overcome limitations of conventional models—namely, heavy reliance on large-scale annotated data and poor adaptability to new users—we propose the Subject-Embedding Network (SE-Net), which jointly models cross-subject discriminability and intra-subject consistency for efficient individual representation learning. Furthermore, we design a conditional scanpath generation model and adopt a fine-tuning-free inference paradigm. Evaluated on multiple eye-tracking datasets, our method achieves high-accuracy personalized predictions using only 1–5 support scanpaths per subject, significantly outperforming state-of-the-art approaches. Experimental results demonstrate strong generalization across unseen users and practical applicability in real-world low-data scenarios.
This study investigates how uncertainty and semantic meaning jointly guide human eye movements in dynamic scenes. We propose the first closed-loop computational model integrating boundary uncertainty estimation with semantic object segmentation: a Bayesian filter recursively models object boundaries and their uncertainties, which serve as novel oculomotor control signals for active visual exploration. Key contributions include: (1) revealing the critical regulatory role of boundary uncertainty in the exploration–exploitation trade-off; and (2) demonstrating that semantic objects constitute the fundamental units of attention, with the model implicitly reproducing high-level oculomotor phenomena such as inhibition-of-return delays. Evaluated on real-world dynamic scene datasets, the model faithfully replicates human free-viewing behavior—quantitatively matching fixation durations, saccade amplitude distributions, and the balanced pattern of object detection, inspection, and revisiting. Moreover, it generalizes to higher-order scanpath statistics not used during model fitting.
This work addresses the key challenge in eye movement modeling: effectively capturing the diversity of human gaze behavior in response to visual stimuli while jointly accounting for continuous trajectory dynamics and discrete scanpath structure. The study proposes, for the first time, a complementary representation that unifies gaze trajectories and scanpaths within a diffusion-based generative framework. By concatenating these representations as additional input channels, the method enables probabilistic generation of human-like fixations without modifying the backbone architecture. Furthermore, a distribution-aware evaluation framework based on Continuous Ranked Probability Score (CRPS) is introduced to better capture the intrinsic variability of gaze behavior. The approach achieves state-of-the-art performance on both task-driven visual search (with target-present and target-absent conditions) and free-viewing benchmarks, demonstrating the efficacy of joint modeling and distribution-aware assessment.
This work addresses a key limitation in existing eye movement scanpath similarity metrics, which predominantly rely on spatial and temporal alignment while neglecting semantic equivalence between fixated regions. To overcome this, the study introduces visual language models (VLMs) to generate context-aware textual descriptions for individual fixations—either via image patches or token-based strategies—and aggregates these into a scanpath-level semantic representation. Semantic similarity is then computed using embedding- and lexicon-based natural language processing metrics. The resulting approach yields an interpretable, content-aware similarity measure that effectively complements traditional geometric alignment methods. Experiments on free-viewing eye-tracking data demonstrate that semantic similarity captures information partially orthogonal to spatial alignment, revealing that scanpaths with divergent spatial distributions can nonetheless exhibit high semantic consistency.
This study addresses the scarcity of real-world eye-tracking data, which is hindered by high annotation costs and privacy concerns, thereby limiting large-scale behavioral modeling research. To overcome this bottleneck, the authors propose an end-to-end synthetic data generation framework that combines authentic iris trajectory extraction with replay in a 3D eye movement simulator, enabling the first large-scale, high-fidelity, and automatically annotated eye-tracking video synthesis. By integrating headless browser automation with trajectory replay techniques, the method effectively circumvents the constraints of real data collection. The released dataset comprises 144 sessions totaling 12 hours of 25-fps synthetic eye-tracking videos, exhibiting highly faithful temporal dynamics (KS D < 0.14), and demonstrates strong validity in script-reading detection tasks.
This study investigates the origins of human eye movement patterns during free viewing and proposes that they emerge as a natural byproduct of optimizing scene understanding under foveal visual constraints. To test this hypothesis, we developed a computational agent equipped with a foveated visual system and trained it—via reinforcement learning or self-supervised strategies—to perform scene understanding tasks without any exposure to human gaze data. Remarkably, the agent spontaneously developed fixation behaviors highly consistent with those of humans, significantly outperforming control models explicitly designed for search or classification tasks. This work provides the first computational modeling evidence establishing an intrinsic link between gaze patterns and perceptual goals, suggesting that human-like fixations arise not from task-specific tuning but from general principles of efficient visual processing under biological constraints.
This study investigates whether multimodal large language models (MLLMs) exhibit human-like visual search under foveated input. Utilizing the COCO-Search18 dataset, we compared human and model gaze trajectories through gaze-contingent foveal simulation and a triaxial evaluation framework. Results indicate that while MLLMs achieve detection efficiency comparable to or exceeding human performance, their fixation patterns are characterized by low entropy and non-sequentiality, lacking the temporal dynamics inherent to human vision. This work reveals how "outcome alignment" can obscure underlying "process heterogeneity," highlighting critical blind spots in conventional metrics for assessing human-like temporal mechanisms. These findings provide essential empirical evidence to guide next-generation research toward achieving genuine cognitive alignment in artificial visual systems.