SocialHumanoid: Towards Expressive Humanoid Behavior via One-Step Co-Speech Motion Generation

πŸ“… 2026-09-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the lack of affective control in humanoid robots and the high latency inherent in joint speech-motion co-generation by proposing a low-latency, highly expressive whole-body motion generation framework based on one-step diffusion models. Methodologically, we construct the AffectMoCap dataset to provide affective supervision and integrate a motion-history conditioning mechanism that enables single-step forward inference for continuous emotional expression. A whole-body controller subsequently translates the generated motions into robot-executable commands online. Experimental results demonstrate that the proposed approach achieves state-of-the-art FrΓ©chet Gesture Distance (FGD) on the BEAT2 benchmark while improving inference speed sixfold. Furthermore, real-world deployment validates its capacity for stable execution during long-horizon tasks.
πŸ“ Abstract
Humanoid robots are increasingly expected to serve as embodied social agents that communicate naturally with humans through face-to-face interaction. During such communication, humanoid robots require body behaviors that are synchronized with speech, affectively expressive, and suitable for real-time execution. However, existing co-speech methods are primarily developed for digital humans and lack joint support for affective control and low-latency continuous generation on physical embodiments. To bridge this gap, we present SocialHumanoid, a system for expressive humanoid behavior via one-step co-speech motion generation. Given response speech and a specified affective condition, SocialHumanoid generates each full-body motion window in a single forward pass and connects successive windows through motion-history conditioning. The generated human motion is further converted online into embodiment-compatible robot references and tracked by a whole-body controller for physical execution. To provide explicit supervision for affective body expression, we further introduce AffectMoCap, a 4-hour dataset captured from two professional actors, containing synchronized speech, body motion, fine-grained hand motion, and emotion annotations. On BEAT2, SocialHumanoid achieves the best FGD among the compared generation methods, competitive speech-motion synchrony, and approximately $6\times$ faster inference than GestureLSM under the same protocol. Perceptual evaluations further show that training with AffectMoCap improves affect recognition from generated body motion, while real-robot experiments demonstrate continuous affect-conditioned behavior and stable long-horizon execution. Our project page is https://rex0191.github.io/SocialHumanoid/.
Problem

Research questions and friction points this paper is trying to address.

Humanoid robots
Co-speech motion generation
Affective expression
Real-time execution
Social agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

Co-Speech Motion Generation
Humanoid Robots
One-Step Generation
Affective Control
AffectMoCap
πŸ”Ž Similar Papers
No similar papers found.