🤖 AI Summary
This study addresses the lack of frame-wise visual guidance in Vision-Language-Action (VLA) model fine-tuning by proposing a gaze prompting mechanism based on eye tracking. Gaze data is collected via virtual reality teleoperation to train a lightweight predictive model that generates gaze prompts, providing robots with fine-grained attentional guidance. During deployment, this approach enables dense supervision without requiring additional hardware. Built upon the π0 architecture and cross-modal alignment techniques, the proposed method improves the average success rate from 26.3% to 56.0% in bimanual manipulation tasks. Furthermore, this work releases GazeMani, an open-source dataset comprising 1,200 trajectories, to facilitate future research in gaze-conditioned robotic manipulation.
📝 Abstract
Vision-Language-Action (VLA) fine-tuning pairs images with actions at every step, yet typically provides only a task-level language instruction, leaving moment-to-moment visual relevance implicit. We introduce \emph{eye-tracker-supervised gaze prompting}, which uses gaze recorded during VR teleoperation to provide frame-level visual guidance for VLA fine-tuning. During training, recorded gaze locations are rendered as crosshairs on the robot's head-camera images. At deployment, a lightweight predictor estimates gaze locations from recent images and the instruction, supplying the same type of visual prompt without an eye tracker or changes to the policy architecture. Instantiated with $π_0$, gaze prompting increases mean success from $26.3\%$ to $56.0\%$ across six real-world bimanual manipulation tasks, with gains also observed when a single policy is trained on all six tasks. We release \textsc{GazeMani}, a dataset of $1{,}200$ teleoperated trajectories with synchronized gaze.