GAE: General Action Expert for Real-Time Humanoid Teleoperation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited diversity of whole-body behaviors and human-robot synchronization latency in humanoid teleoperation by proposing a unified learning framework. Methodologically, we construct a large-scale standardized motion dataset and adopt a two-stage training paradigm that decouples a privileged generator from a deployment executor to synthesize and execute feasible trajectories. To enhance generalization, multi-source data augmentation, curriculum-based domain randomization, and reinforcement learning-based imitation policies are integrated. Furthermore, an adaptive delay prediction and compensation mechanism is designed to ensure real-time control. Experimental results demonstrate that the proposed system enables humanoid platforms, such as the Unitree G1, to fluidly mirror diverse, agile, and expressive human behaviors with minimal latency in real-world settings.
📝 Abstract
Humanoid avatars extend human physical presence beyond the body, enabling people to participate in social, service, and labor activities through remotely operated robots. This requires teleoperation systems capable of realizing diverse and dynamic whole-body behaviors while maintaining responsive human-robot synchronization. We present General Action Expert(GAE), a unified learning framework for general-purpose, low-latency humanoid whole-body teleoperation. To cover diverse human behaviors, GAE builds a large-scale human motion dataset from heterogeneous sources, including videos, animations, and motion capture, followed by standardization and augmentation. GAE then addresses the noise and embodiment mismatch in human motions with a two-stage training paradigm: a privileged generator policy first tracks human motion references in simulation and rolls out feasible humanoid trajectories; a deployable executor policy then learns to track these generated trajectories under curriculum domain randomization. For responsive human-robot synchronization, GAE introduces a latency-conditioned anticipation mechanism that adaptively compensates for end-to-end delay during real-time teleoperation. Simulation and real-world experiments on Unitree G1 and Westlake O1 robots demonstrate that GAE enables humanoids to smoothly mirror diverse, agile, and expressive human behaviors. Project website: https://wangyf0928.github.io/gae-wlrobotics/
Problem

Research questions and friction points this paper is trying to address.

humanoid teleoperation
whole-body control
embodiment mismatch
real-time synchronization
latency compensation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Humanoid Teleoperation
Two-stage Training Paradigm
Latency-conditioned Anticipation
Curriculum Domain Randomization
Whole-body Control
Y
Yuefan Wang
Westlake Robotics, Westlake University
H
Huaicheng Zhou
Westlake Robotics
X
Xiao He
Westlake Robotics
Z
Zhijie He
Westlake Robotics
M
Mingchuan Yang
Westlake Robotics
H
Huayi Zhang
Westlake Robotics
L
Li Chai
Westlake University
J
Jinxin Liu
Westlake Robotics
Donglin Wang
Donglin Wang
Westlake University
Deep Reinforcement LearningMeta LearningRobot Learning