Towards an Extensible Benchmark for Spoken Dialogue with Social Robots
This study addresses the challenges of clarification, interruption, and real-time collaboration in human-robot spoken dialogue by proposing a comprehensive evaluation framework encompassing barge-in handling, embodied signals, and temporal constraints. Methodologically, the authors construct a scalable benchmark and leverage the Retico incremental processing framework to support low-latency interactions, systematically investigating common obstacles in robotic voice interaction and the underlying mechanisms of embodied signals. Ultimately, this work establishes a standardized testing platform that provides a unified evaluation foundation for advancing research on critical elements of human-robot interaction.