Gaze-to-text Generation: Beyond Categorical Decoding of Human Attention

📅 2026-07-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing gaze decoding methods are constrained to predefined categories, limiting their ability to capture the full spectrum of visual intent. This work proposes Gazette, a generative framework that reframes gaze decoding from a classification task into a natural language generation problem, leveraging multimodal large language models to translate eye movement trajectories into free-form textual descriptions of visual targets. By incorporating "think-aloud" transcripts synthetically generated by large language models for instruction tuning, Gazette substantially enhances modeling of goal-directed attention dynamics. Experimental results demonstrate that Gazette achieves state-of-the-art performance across multiple visual tasks, exhibiting strong generalization capabilities and practical utility in real-world applications.
📝 Abstract
We introduce a novel learning problem: decoding gaze into natural language descriptions of human goals across diverse visual tasks. Unlike prior work, which frames gaze decoding as a discriminative task over predefined categories, we formulate it as a generative learning problem: training a model to produce free-form descriptions that capture the rich nuances and open-ended nature of human intentions beyond fixed labels. To this end, we introduce Gazette, the first gaze-to-text decoding framework. Based on multimodal large language models (MLLMs), Gazette learns to decode gaze scanpaths into natural language for goals that may extend beyond categorical labels and require articulation in natural language. To help Gazette filter out individual differences in gaze behavior and learn the goal-specific spatiotemporal dynamics crucial for generating accurate natural language goal descriptions, we propose a novel strategy that leverages the encyclopedic knowledge and reasoning abilities of a large language model to synthesize natural language explanations of goal-directed attentional behavior called think-aloud transcripts. Instruction tuning on these synthetic narratives allows Gazette to achieve state-of-the-art performance in gaze decoding across multiple tasks, demonstrating its generalizability and versatility, thereby enabling gaze to serve as a powerful, non-intrusive cue for inferring human goals and intentions in diverse scenarios.
Problem

Research questions and friction points this paper is trying to address.

gaze decoding
natural language generation
human intention
visual tasks
goal inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

gaze-to-text generation
generative gaze decoding
multimodal large language models
think-aloud transcripts
goal inference