🤖 AI Summary
This study addresses the challenges of clarification, interruption, and real-time collaboration in human-robot spoken dialogue by proposing a comprehensive evaluation framework encompassing barge-in handling, embodied signals, and temporal constraints. Methodologically, the authors construct a scalable benchmark and leverage the Retico incremental processing framework to support low-latency interactions, systematically investigating common obstacles in robotic voice interaction and the underlying mechanisms of embodied signals. Ultimately, this work establishes a standardized testing platform that provides a unified evaluation foundation for advancing research on critical elements of human-robot interaction.
📝 Abstract
Language models provide a plug-and-play interface between humans and robots, but important challenges remain when speech, dialogue, fast interaction, and collaboration are required. We propose a benchmark for the community to use as a way to explore common spoken dialogue artifacts between robots and humans, including requests for clarification, interruptions, embodied signals (e.g., head nods or facial cues), and time constraints. We also explain our vision to extend the benchmark for other aspects of human-robot interaction that are important to the larger research community. To facilitate the benchmark, we further propose using \textit{Retico}, a real-time communication framework that fulfills important technical requirements to enable robots to have spoken dialogue capabilities.