🤖 AI Summary
This study addresses the inability of Vision-Language-Action (VLA) robots to self-identify execution errors by proposing a strategy-agnostic, socially aware gateway that translates spontaneous human reactions into real-time intervention signals for failure detection and recovery. Methodologically, it introduces a novel asynchronous first-event fusion mechanism integrating causal paralinguistic audio, visual reactions, explicit stop phrases, and robot-relevance estimation to enable low-latency local closed-loop control. Experimental evaluations demonstrate an offline recall rate of 54.6% and an online deployment precision of 91.7%, with an average latency below one second. These results confirm that the proposed approach effectively reduces unwarranted interruptions while significantly enhancing manipulation safety during robotic task execution.
📝 Abstract
Vision-language-action (VLA) policies enable diverse robotic manipulation but can fail during execution without recognizing their own errors. Human observers provide complementary signals, as unexpected robot behavior can trigger rapid vocal, facial, or verbal reactions before failure is completed. We introduce SocialVLA, a local, policy-agnostic social perception gateway that converts spontaneous human reactions into runtime intervention signals for VLA manipulation. SocialVLA combines causal paralinguistic audio detection, visual reaction recognition, explicit stop phrases, and robot-relevance estimation. An asynchronous first-event fusion mechanism triggers a VLA hold from the earliest sufficiently confident signal, while a separate speech channel captures verbal corrections for participant-directed continuation, restart, or instruction revision. We evaluate SocialVLA on physical Unitree G1 manipulation using 15 participants, with 238 annotated intervention-worthy episodes and 1.038 h of non-intervention behavior. Frozen offline replay achieves 54.6% recall and 69.5% precision, while unfiltered audio-video fusion reaches 64.3% recall. Relevance estimation reduces false-stop episodes from 100 to 57 and increases precision from 60.5% to 69.8%. In prospective deployment on an unseen 16th participant, the frozen system achieves 59.5% recall and 91.7% precision. Median detector-to-fusion latency is 47.9 ms, VLA-gate-to-physical-hold latency is 336 ms, and reaction-onset-to-hold latency is 1.021 s. These results demonstrate a complete local pathway from spontaneous social reaction to physical VLA interruption and participant-directed recovery.