π€ AI Summary
This work proposes a training-free, semantics-driven robot control method to enhance the naturalness and efficiency of human-robot interaction in assistive care scenarios. By integrating first-person eye tracking with a vision-language foundation model, the system interprets the semantic context of the userβs gaze to infer manipulation intent in real time and autonomously executes tabletop tasks without requiring task-specific training for skill selection or parameterization. The framework combines eye-tracking signals, semantic scene understanding, and modular robot control, demonstrating strong robustness and intuitiveness across diverse tabletop manipulation tasks. It significantly outperforms baseline eye-gaze control approaches that lack semantic reasoning, highlighting its potential for enabling natural and scalable human-robot interaction.
π Abstract
Designing intuitive interfaces for robotic control remains a central challenge in enabling effective human-robot interaction, particularly in assistive care settings. Eye gaze offers a fast, non-intrusive, and intent-rich input modality, making it an attractive channel for conveying user goals. In this work, we present GAMMA (Gaze Assisted Manipulation for Modular Autonomy), a system that leverages ego-centric gaze tracking and a vision-language model to infer user intent and autonomously execute robotic manipulation tasks. By contextualizing gaze fixations within the scene, the system maps visual attention to high-level semantic understanding, enabling skill selection and parameterization without task-specific training. We evaluate GAMMA on a range of table-top manipulation tasks and compare it against baseline gaze-based control without reasoning. Results demonstrate that GAMMA provides robust, intuitive, and generalizable control, highlighting the potential of combining foundation models and gaze for natural and scalable robot autonomy. Project website: https://gamma0.vercel.app/