Gaze Prompts: Temporally Dense Human Attention for Vision-Language-Action Fine-Tuning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of frame-wise visual guidance in Vision-Language-Action (VLA) model fine-tuning by proposing a gaze prompting mechanism based on eye tracking. Gaze data is collected via virtual reality teleoperation to train a lightweight predictive model that generates gaze prompts, providing robots with fine-grained attentional guidance. During deployment, this approach enables dense supervision without requiring additional hardware. Built upon the π0 architecture and cross-modal alignment techniques, the proposed method improves the average success rate from 26.3% to 56.0% in bimanual manipulation tasks. Furthermore, this work releases GazeMani, an open-source dataset comprising 1,200 trajectories, to facilitate future research in gaze-conditioned robotic manipulation.
📝 Abstract
Vision-Language-Action (VLA) fine-tuning pairs images with actions at every step, yet typically provides only a task-level language instruction, leaving moment-to-moment visual relevance implicit. We introduce \emph{eye-tracker-supervised gaze prompting}, which uses gaze recorded during VR teleoperation to provide frame-level visual guidance for VLA fine-tuning. During training, recorded gaze locations are rendered as crosshairs on the robot's head-camera images. At deployment, a lightweight predictor estimates gaze locations from recent images and the instruction, supplying the same type of visual prompt without an eye tracker or changes to the policy architecture. Instantiated with $π_0$, gaze prompting increases mean success from $26.3\%$ to $56.0\%$ across six real-world bimanual manipulation tasks, with gains also observed when a single policy is trained on all six tasks. We release \textsc{GazeMani}, a dataset of $1{,}200$ teleoperated trajectories with synchronized gaze.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
fine-tuning
gaze prompting
visual guidance
bimanual manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gaze Prompting
Vision-Language-Action Fine-Tuning
Eye-tracker Supervision
Lightweight Gaze Predictor
Bimanual Manipulation
🔎 Similar Papers
No similar papers found.
Yihan Zhou
Yihan Zhou
Tsinghua University
ControlRobotics
R
Rui Yan
Department of Automation, Tsinghua University
M
Mingcong Li
LingYu Robotics
Z
Zheyuan Huang
LingYu Robotics
X
Xu Yang
Department of Automation, Tsinghua University
X
Xueyang Guo
Department of Automation, Tsinghua University; Beijing Key Laboratory of Embodied Intelligence Systems; Institute for Embodied Intelligence and Robotics, Tsinghua University
Yilin Mo
Yilin Mo
Associate Professor, Department of Automation, Tsinghua University
Control TheorySignal ProcessingOptimization