prompt-based scanpath generation

Designs, builds, and evaluates models that predict or generate human scanpaths—sequences of fixations and saccades, including reading scanpaths—conditioned on prompts, task instructions, or viewer identity. Implements prompt-based conditioning and related techniques to adapt predicted eye-movement behavior without changing model architecture.

prompt-basedscanpathgeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.46
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Unified Dynamic Scanpath Predictors Outperform Individually Trained Neural Models

May 05, 2024
FA
Fares Abawi
🏛️ University of Hamburg

Existing scanpath prediction methods predominantly rely on population-level models, neglecting inter-individual heterogeneity in eye movements and thus limiting applicability to real-world social human–computer interaction. Method: To address individual gaze prediction in dynamic social videos, we propose the first unified deep learning model that jointly encodes fixation history and social cues via a gated recurrent fusion mechanism, augmented with sequence-wise attention for dynamic saliency representation. The model implicitly learns both universal attention mechanisms and subject-specific patterns within a single architecture. Contribution/Results: Evaluated on a free-viewing social video dataset, it achieves performance on par with or superior to subject-specific models. Large-scale experiments demonstrate that late fusion significantly outperforms early fusion, validating the efficacy of combining universal representations with personalized supervision. Notably, this is the first work to enable cross-subject generalization—predicting diverse observers’ scanpaths using a single model—thereby substantially improving inter-subject transferability.

Improve human-robot interaction via dynamic gaze predictionPredict diverse individual scanpaths in videosUnified model outperforms individual neural models

Decoding Reading Goals from Eye Movements

Oct 28, 2024
OS
Omer Shubi
🏛️ Technion - Israel Institute of Technology | Massachusetts Institute of Technology

This work investigates whether readers’ reading goals—information seeking versus comprehension reading—can be decoded in real time from their eye movement trajectories. To address this, we introduce the first large-scale, annotated eye-tracking dataset and propose a novel Transformer architecture that jointly models scanpath representations and contextualized semantic features from pretrained language models; we further incorporate a mixed-effects model to disentangle text difficulty from inter-subject variability. Methodologically, our approach innovatively integrates sequential visual behavior with linguistic semantics and employs ensemble learning to enhance robustness. Experiments demonstrate that the model achieves high-accuracy goal classification early in reading—well before completion—and identify key determinants of classification difficulty, including textual features (e.g., information density, structural complexity) and individual differences. These findings substantially advance the understanding of oculomotor correlates underlying distinct cognitive mechanisms in the two reading modes.

Decoding reading goals from eye movementsDistinguishing information seeking vs comprehensionReal-time prediction using transformer-based models

This work addresses the limitations of traditional scanpath modeling, which relies on handcrafted architectures and struggles to flexibly incorporate conditional information such as task instructions or individual differences. For the first time, the authors introduce a vision-language model to this task by discretizing gaze coordinates into textual tokens and formulating scanpath prediction as an autoregressive sequence generation problem. Through prompt engineering, the framework jointly models multiple factors—including viewer identity, task goals, and fixation durations—within a unified architecture. The approach not only enables precise computation of information gain per fixation but also offers potential for behavioral intervention. Evaluated on MIT1003, it achieves an information gain of 2.18 bits, a 46% improvement over DeepGaze III, establishes new state-of-the-art results across multiple benchmarks, and successfully reproduces known eye movement phenomena.

oculomotor behaviorscanpath predictionsequence modeling

Few-shot Personalized Scanpath Prediction

Apr 07, 2025
RX
Ruoyu Xue
🏛️ Stony Brook University | EPFL | The University of Adelaide

To address the challenge of personalized scanpath prediction under few-shot settings, this paper introduces the novel Few-Shot Personalized Scanpath Prediction (FS-PSP) task. To overcome limitations of conventional models—namely, heavy reliance on large-scale annotated data and poor adaptability to new users—we propose the Subject-Embedding Network (SE-Net), which jointly models cross-subject discriminability and intra-subject consistency for efficient individual representation learning. Furthermore, we design a conditional scanpath generation model and adopt a fine-tuning-free inference paradigm. Evaluated on multiple eye-tracking datasets, our method achieves high-accuracy personalized predictions using only 1–5 support scanpaths per subject, significantly outperforming state-of-the-art approaches. Experimental results demonstrate strong generalization across unseen users and practical applicability in real-world low-data scenarios.

Address data-intensive training in personalized scanpath predictionCapture unique individual scanpath patterns efficientlyPredict scanpaths for unseen subjects with minimal data

A Robotics-Inspired Scanpath Model Reveals the Importance of Uncertainty and Semantic Object Cues for Gaze Guidance in Dynamic Scenes

Aug 02, 2024
VM
Vito Mengers
🏛️ Technische Universität Berlin | Humboldt-Universität zu Berlin

This study investigates how uncertainty and semantic meaning jointly guide human eye movements in dynamic scenes. We propose the first closed-loop computational model integrating boundary uncertainty estimation with semantic object segmentation: a Bayesian filter recursively models object boundaries and their uncertainties, which serve as novel oculomotor control signals for active visual exploration. Key contributions include: (1) revealing the critical regulatory role of boundary uncertainty in the exploration–exploitation trade-off; and (2) demonstrating that semantic objects constitute the fundamental units of attention, with the model implicitly reproducing high-level oculomotor phenomena such as inhibition-of-return delays. Evaluated on real-world dynamic scene datasets, the model faithfully replicates human free-viewing behavior—quantitatively matching fixation durations, saccade amplitude distributions, and the balanced pattern of object detection, inspection, and revisiting. Moreover, it generalizes to higher-order scanpath statistics not used during model fitting.

Attention AllocationUncertaintyVisual Guidance

Latest Papers

What's happening recently
View more

This work addresses the key challenge in eye movement modeling: effectively capturing the diversity of human gaze behavior in response to visual stimuli while jointly accounting for continuous trajectory dynamics and discrete scanpath structure. The study proposes, for the first time, a complementary representation that unifies gaze trajectories and scanpaths within a diffusion-based generative framework. By concatenating these representations as additional input channels, the method enables probabilistic generation of human-like fixations without modifying the backbone architecture. Furthermore, a distribution-aware evaluation framework based on Continuous Ranked Probability Score (CRPS) is introduced to better capture the intrinsic variability of gaze behavior. The approach achieves state-of-the-art performance on both task-driven visual search (with target-present and target-absent conditions) and free-viewing benchmarks, demonstrating the efficacy of joint modeling and distribution-aware assessment.

eye-tracking trajectorygaze modelinghuman gaze variability

This work addresses a key limitation in existing eye movement scanpath similarity metrics, which predominantly rely on spatial and temporal alignment while neglecting semantic equivalence between fixated regions. To overcome this, the study introduces visual language models (VLMs) to generate context-aware textual descriptions for individual fixations—either via image patches or token-based strategies—and aggregates these into a scanpath-level semantic representation. Semantic similarity is then computed using embedding- and lexicon-based natural language processing metrics. The resulting approach yields an interpretable, content-aware similarity measure that effectively complements traditional geometric alignment methods. Experiments on free-viewing eye-tracking data demonstrate that semantic similarity captures information partially orthogonal to spatial alignment, revealing that scanpaths with divergent spatial distributions can nonetheless exhibit high semantic consistency.

eye-trackingNLP metricsscanpath similarity

This study addresses the scarcity of real-world eye-tracking data, which is hindered by high annotation costs and privacy concerns, thereby limiting large-scale behavioral modeling research. To overcome this bottleneck, the authors propose an end-to-end synthetic data generation framework that combines authentic iris trajectory extraction with replay in a 3D eye movement simulator, enabling the first large-scale, high-fidelity, and automatically annotated eye-tracking video synthesis. By integrating headless browser automation with trajectory replay techniques, the method effectively circumvents the constraints of real data collection. The released dataset comprises 144 sessions totaling 12 hours of 25-fps synthetic eye-tracking videos, exhibiting highly faithful temporal dynamics (KS D < 0.14), and demonstrates strong validity in script-reading detection tasks.

behavioral modelingeye movementscript reading detection

This study investigates the origins of human eye movement patterns during free viewing and proposes that they emerge as a natural byproduct of optimizing scene understanding under foveal visual constraints. To test this hypothesis, we developed a computational agent equipped with a foveated visual system and trained it—via reinforcement learning or self-supervised strategies—to perform scene understanding tasks without any exposure to human gaze data. Remarkably, the agent spontaneously developed fixation behaviors highly consistent with those of humans, significantly outperforming control models explicitly designed for search or classification tasks. This work provides the first computational modeling evidence establishing an intrinsic link between gaze patterns and perceptual goals, suggesting that human-like fixations arise not from task-specific tuning but from general principles of efficient visual processing under biological constraints.

foveated visionfree-viewinghuman fixation patterns

This study investigates whether multimodal large language models (MLLMs) exhibit human-like visual search under foveated input. Utilizing the COCO-Search18 dataset, we compared human and model gaze trajectories through gaze-contingent foveal simulation and a triaxial evaluation framework. Results indicate that while MLLMs achieve detection efficiency comparable to or exceeding human performance, their fixation patterns are characterized by low entropy and non-sequentiality, lacking the temporal dynamics inherent to human vision. This work reveals how "outcome alignment" can obscure underlying "process heterogeneity," highlighting critical blind spots in conventional metrics for assessing human-like temporal mechanisms. These findings provide essential empirical evidence to guide next-generation research toward achieving genuine cognitive alignment in artificial visual systems.

Attention AlignmentFoveated VisionHuman Visual Search

Hot Scholars

PM

Plinio Moreno

Institute for Systems and Robotics, Instituto Superior Tecnico (ISR/IST), LARSyS, Univ Lisboa
object manipulationcomputer visionroboticsmachine learning
AB

Alexandre Bernardino

Institute for Systems and Robotics (ISR/IST), LARSyS, Instituto Superior Técnico, Univ Lisboa
Computer VisionRobotics
AC

Aladine Chetouani

Institut Galilée - L2TI - Multimedia Team
Image Quality AssessmentVideo AnalysisDepp LearningPattern Recognition
AB

Alessandro Bruno

Associate Professor of Computer Science at IULM University - Principal Investigator
Image ProcessingPerceptionComputer VisionBiomedical Imaging