🤖 AI Summary
This study addresses the challenges of jitter and implicit trajectory learning in gaze estimation under natural head-eye movements by proposing a causal multi-frame gaze estimation framework. The method incorporates explicit first-order gaze priors as compact kinematic tokens fed back into the model, integrating facial and ocular visual evidence for joint estimation. Furthermore, it introduces differential gaze trajectory tokens that exploit translation invariance to extract subject-agnostic motion features, thereby eliminating systematic saccadic biases. Efficient modeling is achieved through a causal Transformer decoder with cross-attention mechanisms. Experimental results demonstrate that this framework reduces the mean angular error by approximately 1.0° on the Gaze360 dataset and attains state-of-the-art baseline performance on the EVE dataset.
📝 Abstract
Gaze estimation under natural head-eye motion underpins applications from driver monitoring to human-computer interaction. Single-frame methods predict each frame independently, so consecutive outputs fluctuate as jitter. Multi-frame methods reduce this, but they learn motion implicitly inside appearance features, so the gaze trajectory is never an explicit variable. We propose EyeTAG (Eye Trajectory-Aware Gaze Estimation), a causal multi-frame framework built around an explicit first-order gaze prior: at each step it differentiates its own recent predictions and feeds the resulting trajectory back as a compact kinematic token. Because differencing is translation-invariant in gaze space, this token carries subject-invariant motion rather than personal gaze offsets. Face and eye streams supply visual evidence, fused by cross-attention and a causal Transformer decoder. EyeTAG reduces the mean angular error by about 1.0$^\circ$ on Gaze360 and performs on par with the strongest baseline on EVE (2.56$^\circ$ vs. 2.58$^\circ$). Within-model ablations, which keep the encoder and the rest of the architecture fixed and vary only the gaze history, show that the differential formulation, rather than temporal context alone, removes the systematic saccade bias that persists even with an absolute gaze-history prior. Our code is available at https://github.com/peter8366/EyeTAG.