BiGym 2.0: Benchmarking Learned and Agent-Developed Policies for Humanoid Household Manipulation

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of a unified benchmark for whole-body coordinated household manipulation tasks in humanoid robots. Based on the Unitree G1 platform, we construct a simulation environment comprising 20 tasks and provide corresponding VR demonstration data. This work presents the first fair evaluation of diverse policy learning methods under a whole-body controller, encompassing Vision-Language-Action (VLA) fine-tuning, imitation learning, demonstration-driven reinforcement learning, and cold-start coding agents. Experimental results demonstrate that VLA approaches achieve the best average performance, while coding agents significantly outperform traditional reinforcement learning baselines in bimanual tasks; however, complex scenarios such as multi-object transportation remain challenging. All environments and datasets have been fully open-sourced to facilitate future research.
📝 Abstract
Humanoid household manipulation requires the arms to act while the body balances, steps and changes posture. We present BiGym 2.0, an adaptation of BiGym for the Unitree G1 across 20 household tasks using a unified whole-body controller for demonstration and evaluation. The suite provides 60 native human virtual-reality demonstrations per task with synchronised multi-camera views and full-body execution records. We benchmark vision-language-action fine-tuning, imitation learning, demo-driven reinforcement learning, and cold-start coding agents given the interaction budget of online reinforcement learning. With the same onboard views, proprioception and whole-body controller for every method, vision-language-action fine-tuning has the highest nine-task mean, and agent-developed programs outperform every demo-driven reinforcement learning baseline on this mean and lead on bimanual reaching. Cross-workspace stacking remains open, $π_{0.5}$ stays low on pick-box, and multi-object transport is hard for imitation learning, demo-driven reinforcement learning and coding agents. All environments, human demonstrations, and evaluation traces are open-sourced at https://github.com/swirl-uk/BiGym2.
Problem

Research questions and friction points this paper is trying to address.

Humanoid Manipulation
Whole-body Control
Benchmarking
Household Tasks
Policy Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Humanoid Manipulation
Whole-body Control
Vision-Language-Action
Benchmark
Coding Agents
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zexi Zhang
Department of Computing, Imperial College London, UK
Z
Zecheng Zhu
Department of Computing, Imperial College London, UK
Z
Zidong Chen
Department of Computing, Imperial College London, UK
Z
Zulkhuu Tuya
Department of Computing, Imperial College London, UK
Stephen James
Stephen James
Robot Learning Lab
Robot LearningRoboticsMachine LearningAI