ECHO-G: Embodied Co-speech Humanoid mOtion Generation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of coordinating speech prosody, linguistic content, and embodied motion in whole-body co-speech gesture generation for humanoid robots. To this end, it proposes the ECHO-G framework, which pioneers modeling many-to-one mappings directly within the robot space to replace conventional retargeting pipelines. Furthermore, a Speech-Grounded Diffusion Transformer (SGDiT) is designed and integrated with rectified flow matching to achieve joint, fine-grained conditional control over both audio and text modalities. This work also establishes a novel dataset and evaluation benchmark. Experimental results demonstrate that the proposed approach significantly outperforms existing baselines in generation quality and efficiency, validating the advantages of joint conditioning. Finally, the method is successfully deployed on a physical humanoid robot.
📝 Abstract
Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly conditioned on speech audio and timed transcripts. Its Speech-Grounded Diffusion Transformer (SGDiT) combines frame-aligned acoustic features with token-level linguistic context, preserving their distinct granularities. Trained with rectified flow matching, it models one-to-many utterance-motion relationships directly in robot space. To support training and evaluation, we introduce a BEAT2-derived audio-text-robot dataset and a benchmark covering co-speech characteristics, robot-motion quality, and runtime efficiency. Comparative evaluation supports direct robot-space generation over the evaluated human-motion generation and retargeting pipelines, while modality ablations highlight the benefits of joint audio-text conditioning. We further demonstrate deployment on a physical humanoid robot. A complementary video-rating study also favors joint conditioning over the alternatives. The dataset and training, inference, and evaluation code are available through our project page.
Problem

Research questions and friction points this paper is trying to address.

co-speech motion
humanoid robots
motion generation
speech prosody
Innovation

Methods, ideas, or system contributions that make the work stand out.

Co-speech Motion Generation
Diffusion Transformer
Rectified Flow Matching
Humanoid Robot
Joint Audio-Text Conditioning
🔎 Similar Papers
No similar papers found.
Y
Yizhao Li
Beihang University; Mondo Robotics
P
Pusen Gao
The Hong Kong University of Science and Technology; Mondo Robotics
M
Ming Wang
Beihang University
Shaojie Shen
Shaojie Shen
Associate Professor, Hong Kong University of Science and Technology
Robotics
S
Shuo Yang
Mondo Robotics
H
Hao Xu
Nanjing University