🤖 AI Summary
This study systematically compares LSTM, Transformer, and a novel joint classification-prediction (ClPr) network for short-term gaze prediction, with emphasis on inter-subject heterogeneity and analysis of typical versus extreme error cases. Method: We employ a three-layer LSTM, a Transformer encoder, and the ClPr architecture—designed to jointly model eye movement event classification and gaze position prediction—within a unified framework. Contribution/Results: We introduce the first dedicated analytical framework for high-percentile prediction errors, revealing substantial divergence between median and long-tail error distributions across individuals—highlighting the necessity of personalized modeling. Experimental results show that LSTM achieves the highest overall robustness; ClPr and Transformer outperform LSTM in post-saccadic prediction accuracy; and model performance exhibits task-dependent trade-offs across distinct eye movement types (e.g., fixations, saccades). These findings establish a new evaluation paradigm and optimization pathway for individualized gaze prediction.
📝 Abstract
Gaze prediction is a diverse field of study with multiple research focuses and practical applications. This article investigates how recurrent neural networks and transformers perform short-term gaze prediction. We used three models: a three-layer long-short-term memory (LSTM) network, a simple transformer-encoder model (TF), and a classification-predictor network (ClPr), which simultaneously classifies the signal into eye movement events and predicts the positions of gaze. The performance of the models was evaluated for ocular fixations and saccades of various amplitudes and as a function of individual differences in both typical and extreme cases. On average, LSTM performed better on fixations and saccades, whereas TF and ClPr demonstrated more precise results for post-saccadic periods. In extreme cases, the best-performing models vary depending on the type of eye movement. We reviewed the difference between the median $P_{50}$ and high-percentile $P_{95}$ error profiles across subjects. The subjects for which the models perform the best overall do not necessarily exhibit the lowest $P_{95}$ values, which supports the idea of analyzing extreme cases separately in future work. We explore the trade-offs between the proposed solutions and provide practical insights into model selection for gaze prediction.