🤖 AI Summary
This study addresses the scalability limitations of conventional approaches for humanoid mobile manipulation, which rely heavily on task-specific reward engineering and demonstration collection. To overcome these challenges, this work proposes a hierarchical control framework that pioneers the use of large language models to directly generate executable high-level policy code. By incorporating closed-loop feedback based on trajectories and video frames, the method enables self-correction and iterative optimization, thereby entirely eliminating the need for manual reward design and demonstration data. Experimental results demonstrate that the proposed approach outperforms reinforcement learning baselines in simulation while approaching demonstration-level performance. Furthermore, it achieves successful zero-shot sim-to-real transfer with additional performance gains on physical hardware, exhibiting strong generalizability and scalability.
📝 Abstract
For humanoids to be useful in everyday environments, they must perform a wide range of tasks that couple locomotion and manipulation. Existing approaches commonly acquire a loco-manipulation policy through reward engineering or demonstrations followed by task-specific training, making it costly to scale to new tasks. In this work, we propose a hierarchical approach to humanoid loco-manipulation that eliminates these per-task requirements. HuGo, Humanoid policy code Generation, uses a Large Language Model (LLM) to generate executable, closed-loop high-level policy code from a task description on top of a frozen low-level whole-body policy. Given the task, observation, and command specifications, the LLM constructs the task logic in code. HuGo then refines the policy from its rollouts using numerical trajectories and selected video frames to produce feedback and targeted code updates. Across five simulation tasks, using two different low-level policies, HuGo substantially outperforms a high-level reinforcement learning baseline and approaches the performance of a demonstration-based baseline. We achieve this level of performance without task-specific reward design or demonstration collection. We further demonstrate zero-shot transfer of simulation-generated policies to hardware and show that applying the same refinement loop to real-world rollouts can further improve transfer performance without expert demonstrations or policy retraining. Project website is https://iconlab.negarmehr.com/HuGo/