🤖 AI Summary
This study addresses the challenge of coordinating speech prosody, linguistic content, and embodied motion in whole-body co-speech gesture generation for humanoid robots. To this end, it proposes the ECHO-G framework, which pioneers modeling many-to-one mappings directly within the robot space to replace conventional retargeting pipelines. Furthermore, a Speech-Grounded Diffusion Transformer (SGDiT) is designed and integrated with rectified flow matching to achieve joint, fine-grained conditional control over both audio and text modalities. This work also establishes a novel dataset and evaluation benchmark. Experimental results demonstrate that the proposed approach significantly outperforms existing baselines in generation quality and efficiency, validating the advantages of joint conditioning. Finally, the method is successfully deployed on a physical humanoid robot.
📝 Abstract
Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly conditioned on speech audio and timed transcripts. Its Speech-Grounded Diffusion Transformer (SGDiT) combines frame-aligned acoustic features with token-level linguistic context, preserving their distinct granularities. Trained with rectified flow matching, it models one-to-many utterance-motion relationships directly in robot space. To support training and evaluation, we introduce a BEAT2-derived audio-text-robot dataset and a benchmark covering co-speech characteristics, robot-motion quality, and runtime efficiency. Comparative evaluation supports direct robot-space generation over the evaluated human-motion generation and retargeting pipelines, while modality ablations highlight the benefits of joint audio-text conditioning. We further demonstrate deployment on a physical humanoid robot. A complementary video-rating study also favors joint conditioning over the alternatives. The dataset and training, inference, and evaluation code are available through our project page.