๐ค AI Summary
This study addresses the challenge of degraded situational awareness among human operators in high-risk environments, often caused by delayed or unclear access to robot-identified hazards. To bridge this gap, the authors propose an integrated human-robot interaction framework that combines vision-language models (VLMs) with immersive virtual reality (VR). In this approach, a robot autonomously explores a virtual environment and leverages a VLM to detect and semantically annotate risks, which are then intuitively conveyed to the operator through a VR interface enriched with contextual labels. This work represents the first unified system that couples VLM-driven risk perception with human-centered safety communication. User studies demonstrate that operators significantly prefer the annotated VR interface over baseline alternatives, rating it higher in clarity, usefulness, and comfortโthereby validating its efficacy in enhancing situational awareness.
๐ Abstract
In high-risk environments such as disaster response, situational awareness depends not only on detecting hazards but also on communicating them clearly to human operators. Vision Language Models (VLMs) have shown strong potential for scene understanding in safety-critical settings, yet their value as part of human-facing robotic systems remains underexplored. We present a VR-based Human Robot Interaction framework for studying how VLM-assisted robots can support situational awareness in simulated hazardous environments. In our system, a robot explores a virtual scene and queries a VLM to identify potential hazards and annotate user-facing points of interest. These annotations are presented to a human operator through an immersive VR interface. This framework enables controlled evaluation of both robotic hazard identification and the communication of safety-critical information to users. Results from our study indicate that the annotated VR interface was preferred over the unannotated baseline and that participants reported high clarity, usefulness, and comfort when interacting with the system. These findings suggest that combining VLM-based robotic perception with immersive visualization is a promising approach for supporting situational awareness in hazardous settings.