π€ AI Summary
This study addresses the lack of affective control in humanoid robots and the high latency inherent in joint speech-motion co-generation by proposing a low-latency, highly expressive whole-body motion generation framework based on one-step diffusion models. Methodologically, we construct the AffectMoCap dataset to provide affective supervision and integrate a motion-history conditioning mechanism that enables single-step forward inference for continuous emotional expression. A whole-body controller subsequently translates the generated motions into robot-executable commands online. Experimental results demonstrate that the proposed approach achieves state-of-the-art FrΓ©chet Gesture Distance (FGD) on the BEAT2 benchmark while improving inference speed sixfold. Furthermore, real-world deployment validates its capacity for stable execution during long-horizon tasks.
π Abstract
Humanoid robots are increasingly expected to serve as embodied social agents that communicate naturally with humans through face-to-face interaction. During such communication, humanoid robots require body behaviors that are synchronized with speech, affectively expressive, and suitable for real-time execution. However, existing co-speech methods are primarily developed for digital humans and lack joint support for affective control and low-latency continuous generation on physical embodiments. To bridge this gap, we present SocialHumanoid, a system for expressive humanoid behavior via one-step co-speech motion generation. Given response speech and a specified affective condition, SocialHumanoid generates each full-body motion window in a single forward pass and connects successive windows through motion-history conditioning. The generated human motion is further converted online into embodiment-compatible robot references and tracked by a whole-body controller for physical execution. To provide explicit supervision for affective body expression, we further introduce AffectMoCap, a 4-hour dataset captured from two professional actors, containing synchronized speech, body motion, fine-grained hand motion, and emotion annotations. On BEAT2, SocialHumanoid achieves the best FGD among the compared generation methods, competitive speech-motion synchrony, and approximately $6\times$ faster inference than GestureLSM under the same protocol. Perceptual evaluations further show that training with AffectMoCap improves affect recognition from generated body motion, while real-robot experiments demonstrate continuous affect-conditioned behavior and stable long-horizon execution. Our project page is https://rex0191.github.io/SocialHumanoid/.