π€ AI Summary
This work addresses the challenge of detecting human-initiated interaction in domestic environments without relying on keyword triggers. The authors propose a multimodal perception framework that fuses audio and visual cues, integrating non-verbal signals such as sound source localization, head orientation, and gaze duration. By leveraging coordinated multi-camera sensing and a state-transition model, the approach uniquely incorporates gaze behavior and temporal dynamics into a multisensor fusion mechanism. Implemented on a mobile robotic platform using the Robot Operating System (ROS), the system demonstrates significantly improved accuracy and robustness in detecting interaction intent, enabling more natural and unobtrusive initiation of humanβrobot interaction.
π Abstract
This paper describes an initiation of interaction(IoI) detection framework without keywords for human-robot interaction(HRI) based on audio and vision sensor fusion in a domestic environment. In the proposed framework, the robot has its own audio and vision sensors, and can employ external vision sensor for stable human detection and tracking. When the user starts to speak while looking at the robot, the robot can localize his or her position by its sound source localization together with human tracking information. Then the robot can detect the IoI if it perceives the face of the speaker faces the robot. In case that the user does not speak directly, the robot can also detect the IoI if he or she looks at the robot for more than predefined periods of time. A state transition model for the proposed IoI detection framework is designed and verified by experiments with a mobile robot. In order to implement and associate our model in a robot architecture, all the components are implemented and integrated in the Robot Operating System(ROS) environment.