🤖 AI Summary
This study addresses the challenge of effectively transferring large-scale human motion priors to perception-driven object interaction tasks for humanoid robots operating in unstructured environments. To this end, the authors extend the Generative Pre-trained Controller (GPC) framework to construct scene-affordance-based expert interaction policies. Through knowledge distillation with dual complementary objectives, privileged expert knowledge is transferred into a student policy that relies solely on onboard perception, thereby enabling multimodal fusion and whole-body control. The effectiveness of the proposed approach is validated across diverse contact-rich, whole-body interaction tasks. Ultimately, this work provides an efficient adaptation strategy for deploying generative motion priors in real-world settings, bridging the gap between large-scale pretraining and practical perceptual constraints in humanoid robotics.
📝 Abstract
Humanoid robots operating in unstructured environments must combine robust whole-body control with the ability to perceive and physically interact with surrounding objects. While large-scale human motion data provides powerful priors for natural and versatile humanoid control, effectively transferring such priors to perception-driven object interaction remains challenging. To address this bottleneck, we propose a framework that extends the recently proposed Generative Pretrained Controller (GPC) from general human motion to full-body humanoid-environment interaction. First, we adapt GPC into interaction experts conditioned on scene affordance cues and privileged state information. These experts leverage the pretrained human motion prior while learning task-specific contact behaviors, including reaching toward objects, grasping environmental supports for stabilization, and pushing movable objects. Second, we introduce a perception-driven student that retains the pretrained GPC policy and distills interaction skills from the experts using onboard sensory observations. To bridge the gap between privileged expert observations and sensory inputs, we propose two complementary training objectives that enable effective adaptation of the pretrained motion prior during distillation. Notably, our experiments across multiple whole-body interaction tasks demonstrate that large-scale generative human motion priors provide an effective foundation for learning deployable policies for humanoid interactions in contact-rich real-world environments.