Intent at a Glance: Gaze-Guided Robotic Manipulation via Foundation Models

πŸ“… 2026-01-08
πŸ›οΈ arXiv.org
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work proposes a training-free, semantics-driven robot control method to enhance the naturalness and efficiency of human-robot interaction in assistive care scenarios. By integrating first-person eye tracking with a vision-language foundation model, the system interprets the semantic context of the user’s gaze to infer manipulation intent in real time and autonomously executes tabletop tasks without requiring task-specific training for skill selection or parameterization. The framework combines eye-tracking signals, semantic scene understanding, and modular robot control, demonstrating strong robustness and intuitiveness across diverse tabletop manipulation tasks. It significantly outperforms baseline eye-gaze control approaches that lack semantic reasoning, highlighting its potential for enabling natural and scalable human-robot interaction.

Technology Category

Intelligent Robots: ManipulationHumans and AI: Human-Aware Planning and Behavior PredictionComputer Vision: Language and Vision

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsEconomics, Online Markets and Human Computation: LLM based quality controls for crowd workResponsible Web: Machine-in-the-loop, human agency and autonomy
πŸ“ Abstract
Designing intuitive interfaces for robotic control remains a central challenge in enabling effective human-robot interaction, particularly in assistive care settings. Eye gaze offers a fast, non-intrusive, and intent-rich input modality, making it an attractive channel for conveying user goals. In this work, we present GAMMA (Gaze Assisted Manipulation for Modular Autonomy), a system that leverages ego-centric gaze tracking and a vision-language model to infer user intent and autonomously execute robotic manipulation tasks. By contextualizing gaze fixations within the scene, the system maps visual attention to high-level semantic understanding, enabling skill selection and parameterization without task-specific training. We evaluate GAMMA on a range of table-top manipulation tasks and compare it against baseline gaze-based control without reasoning. Results demonstrate that GAMMA provides robust, intuitive, and generalizable control, highlighting the potential of combining foundation models and gaze for natural and scalable robot autonomy. Project website: https://gamma0.vercel.app/
Problem

Research questions and friction points this paper is trying to address.

human-robot interaction
gaze-based control
intent inference
assistive robotics
intuitive interfaces
Innovation

Methods, ideas, or system contributions that make the work stand out.

gaze-guided manipulation
foundation models
vision-language model
human-robot interaction
intent inference
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
T
Tracey Yee Hsin Tay
University of California, Los Angeles
X
Xu Yan
University of California, Los Angeles
J
Jonathan Ouyang
University of California, Los Angeles
D
Daniel Wu
University of California, Los Angeles
W
William Jiang
University of California, Los Angeles
Jonathan Kao
Jonathan Kao
University of California, Los Angeles
NeuroscienceNeural EngineeringStatistical Signal ProcessingMachine Learning
Y
Yuchen Cui
University of California, Los Angeles