🤖 AI Summary
This study addresses the frequent failures in human-robot interaction caused by inconsistencies between humans' implicit assumptions and the robot's world state, which traditional methods handle only through post-hoc recovery. We propose an IoT-augmented framework that uniformly models such failures as "world state mismatches." By leveraging large language models to explicate human implicit assumptions and map them to specific mismatch types, and integrating multimodal perception with world models, our approach enables proactive detection and norm-guided repair during the intention formation stage. This work is the first to formally define HRI failures as explicit state mismatches, driving a paradigm shift from passive execution-phase recovery to proactive intention-phase prevention while incorporating social norms and safety constraints. Experiments across ten everyday scenarios demonstrate reliable detection of both visual and latent state mismatches, significantly improving early failure discovery rates.
📝 Abstract
Human-robot interaction (HRI) failures remain a major barrier to deploying robots in real-world environments. Prior work often treats failures as isolated technical faults or focuses on post-hoc recovery behaviors. In practice, many breakdowns arise because humans and robots operate under inconsistent assumptions about the current world state. We propose WSM-Aware HRI, an IoT-enhanced modular framework that unifies diverse HRI breakdowns as World-State Mismatches (WSMs) between a human's instruction-implied assumptions and a robot's grounded world model built from multimodal perception and digital augmentation. A Large Language Model (LLM) is used to make implicit assumptions explicit, map them to a small set of mismatch types, and specify the evidence needed for verification against the robot's world state. WSM-Aware HRI shifts failure handling from execution-time recovery to proactive mismatch detection during intention formation, enabling interventions guided by safety, norm compliance, and multi-user coordination with transparent explanations. We evaluate mismatch identification in ten everyday cases spanning both visual and latent-state mismatches. The system can accurately produce the expected output results, and ablations show that reliable identification depends on appropriate grounding representations and verification-oriented refinement. These results indicate that treating interaction breakdowns as explicit world-state mismatches enables earlier detection of impending failures and offers a principled mechanism for integrating external evidence and social constraints into human-robot interaction.