ViLoMan: Learning Visual-Proprioceptive Whole-Body Loco-Manipulation Skills for Humanoid Robots

📅 2026-09-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
ViLoMan通过学习人类物体互动的物理可执行机器人轨迹,并利用这些轨迹训练统一策略,解决了人形机器人自主全身协调操作难题。
📝 Abstract
Humanoid loco-manipulation requires adaptive whole-body coordination to seamlessly integrate locomotion and physical interaction. Despite recent advances, learning autonomous loco-manipulation remains challenging due to the scarcity of diverse, physically executable robot-object interaction data and the difficulty of learning unified whole-body control directly from onboard observations. We present ViLoMan, a scalable framework for autonomous humanoid loco-manipulation. ViLoMan first transforms partial kinematic demonstrations of human-object interactions into complete, physically executable robot trajectories. It then leverages these trajectories within a teacher-student distillation framework to learn a unified policy that maps egocentric depth observations and proprioceptive measurements directly to joint-level whole-body actions. During deployment, the policy requires neither reference motions nor intermediate commands. We evaluate ViLoMan on door-closing tasks across diverse door configurations and robot initial conditions in both simulation and the real world. Experimental results demonstrate that a single policy enables a Unitree G1 humanoid to complete the full task using only onboard depth sensing and proprioception, while generalizing robustly across task variations and transferring effectively from simulation to reality. Project page: viloman-anonymous.pages.dev.
Problem

Research questions and friction points this paper is trying to address.

humanoid loco-manipulation
whole-body coordination
physically executable data
autonomous learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

autonomous loco-manipulation
teacher-student distillation framework
egocentric depth observations
proprioceptive measurements
whole-body actions
💼 Related Jobs
No related jobs found.
Z
Zejie Tian
State Key Laboratory of AI Safety, Institute of Computing Technology, CAS, China; University of Chinese Academy of Sciences (CAS), China; Beijing Academy of Artificial Intelligence (BAAI)
Ruibing Hou
Ruibing Hou
Institute of Computing Technology, Chinese Academy of Sciences
Computer VisionDeep Learning
B
Bingpeng Ma
State Key Laboratory of AI Safety, Institute of Computing Technology, CAS, China; University of Chinese Academy of Sciences (CAS), China
Börje F. Karlsson
Börje F. Karlsson
Beijing Academy of Artificial Intelligence (BAAI)
Machine Learning SystemsIntelligent AgentsKnowledge MiningMobile ComputingMultilinguality
Shiguang Shan
Shiguang Shan
Professor of Institute of Computing Technology, Chinese Academy of Sciences
Computer VisionPattern RecognitionMachine LearningFace Recognition