Beyond Activation: Gaze Invocation with Visible Status for an Embodied AR Assistant in Co-Located Collaboration

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses inadvertent voice assistant activations and addressee ambiguity in collaborative augmented reality (AR) by proposing a channel management paradigm that integrates gaze-based addressing with embodied state feedback. Leveraging eye tracking for gaze invocation, combined with AR embodied rendering and dynamic hold strategies, the approach systematically governs the activation, maintenance, and release of interaction channels while revealing perceptual blind spots in state awareness following visual attention shifts. A 25-participant user study demonstrates that 98.8% of voice commands were issued after gaze acquisition, effectively ensuring precise interaction channel establishment. However, the findings also expose significant design challenges associated with exit notification mechanisms.
📝 Abstract
Wake words face an inherent trade-off: higher sensitivity reduces missed commands but increases accidental activations. In co-located augmented reality (AR), the system must also determine whether the user is speaking to the assistant or to a nearby person. We introduce a new perspective on this problem: combining gaze-based address with an embodied assistant, visible listening status, and turn management across activation, continued interaction, and release. In our implementation, users look at an assistant anchored in the scene and see activation progress and listening status on its body. Sustained gaze opens a local interaction channel; speech and playback keep it open when attention returns to the task; and inactivity closes it. We evaluated this design with 25 participants. After selecting the assistant's placement and dwell duration, each participant worked with a partner to plan a trip while using the assistant. The study recorded 500 interaction outcomes, including 496 assistant-directed utterance attempts. Gaze acquisition completed before speech for 490 of these attempts (98.8\%): 467 began while the channel remained open, whereas 23 began after it had been released. The other 6 attempts began before acquisition completed. Among 102 reviewed attempts in which gaze left after acquisition but before speech, the configured retention policy kept the channel open for 79 and released it before 23. The results show that an embodied target with visible status can support clear entry into an assistant interaction while revealing a different problem at release: after users look back to their work, they may not see that the assistant has stopped listening. We contribute the implemented gaze-invocation design and design implications for communicating assistant state after visual attention moves elsewhere.
Problem

Research questions and friction points this paper is trying to address.

wake word trade-off
co-located AR collaboration
addressee identification
interaction state awareness
embodied AR assistant
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gaze Invocation
Embodied AR Assistant
Visible Listening Status
Turn Management
Co-Located Collaboration
C
Chenrui Ma
Graduate School of Interdisciplinary Information Studies, The University of Tokyo, Tokyo, Japan
Yoshio Ishiguro
Yoshio Ishiguro
The University of Tokyo
Interaction
Q
Qing Zhang
Interfaculty Initiative in Information Studies, The University of Tokyo, Tokyo, Japan