🤖 AI Summary
This study addresses the challenge of first-person gaze prediction, where inherent human gaze stochasticity complicates the simultaneous modeling of temporal dynamics, task intentions, and visual saliency constraints. To overcome this, the proposed method formulates gaze as a conditional joint distribution and introduces a Conditional Flow Matching (CFM) framework to directly generate spatiotemporal gaze trajectories. By leveraging a video encoder to extract bottom-up visual features alongside a global query mechanism for top-down task information, the approach achieves deep fusion of multi-source representations. This work demonstrates that the proposed framework effectively captures the structured temporal dynamics underlying gaze stochasticity. Extensive evaluations on standard benchmarks reveal state-of-the-art performance, with the generated gaze trajectories exhibiting superior alignment with authentic human behavioral patterns.
📝 Abstract
Egocentric gaze prediction enables many downstream applications but remains challenging, as human gaze is inherently stochastic. This stochasticity is constrained by structured temporal dynamics alternating between fixations and saccades, top-down influences from tasks, and bottom-up visual saliency. Based on this observation, we introduce GazeFlow, a framework that directly models gaze as a joint distribution of temporal gaze positions conditioned upon both top-down and bottom-up information. In particular, GazeFlow uses conditional flow matching (CFM): a learned velocity field iteratively transports a Gaussian noise sample into a plausible gaze trajectory drawn from this joint distribution. The velocity field is conditioned on bottom-up visual features extracted by a video encoder and on top-down task information obtained by globally querying these features. On standard datasets, GazeFlow achieves state-of-the-art performance on per-frame metrics, and the generated trajectories align better with human gaze temporal dynamics.