🤖 AI Summary
This study addresses the lack of a unified benchmark for whole-body coordinated household manipulation tasks in humanoid robots. Based on the Unitree G1 platform, we construct a simulation environment comprising 20 tasks and provide corresponding VR demonstration data. This work presents the first fair evaluation of diverse policy learning methods under a whole-body controller, encompassing Vision-Language-Action (VLA) fine-tuning, imitation learning, demonstration-driven reinforcement learning, and cold-start coding agents. Experimental results demonstrate that VLA approaches achieve the best average performance, while coding agents significantly outperform traditional reinforcement learning baselines in bimanual tasks; however, complex scenarios such as multi-object transportation remain challenging. All environments and datasets have been fully open-sourced to facilitate future research.
📝 Abstract
Humanoid household manipulation requires the arms to act while the body balances, steps and changes posture. We present BiGym 2.0, an adaptation of BiGym for the Unitree G1 across 20 household tasks using a unified whole-body controller for demonstration and evaluation. The suite provides 60 native human virtual-reality demonstrations per task with synchronised multi-camera views and full-body execution records. We benchmark vision-language-action fine-tuning, imitation learning, demo-driven reinforcement learning, and cold-start coding agents given the interaction budget of online reinforcement learning. With the same onboard views, proprioception and whole-body controller for every method, vision-language-action fine-tuning has the highest nine-task mean, and agent-developed programs outperform every demo-driven reinforcement learning baseline on this mean and lead on bimanual reaching. Cross-workspace stacking remains open, $π_{0.5}$ stays low on pick-box, and multi-object transport is hard for imitation learning, demo-driven reinforcement learning and coding agents. All environments, human demonstrations, and evaluation traces are open-sourced at https://github.com/swirl-uk/BiGym2.