🤖 AI Summary
This study addresses the lack of joint evaluation for tool selection, manipulation, and locomotion execution in existing humanoid robot benchmarks. To this end, we construct an 18-task benchmark spanning diverse scenarios alongside a 3.1k demonstration dataset, collecting data through both simulation and real-world experiments on the Unitree G1 platform. Furthermore, we employ the GR00T N1.7 model for probing analysis. This work presents the first closed-loop evaluation of humanoid robots from tool selection to mobile execution, revealing a significant performance gap between these capabilities. Experimental results demonstrate that current policies exhibit low tool selection accuracy and remain highly susceptible to interference from irrelevant instructions. These findings expose critical limitations in existing methodologies and provide clear directions for future research.
📝 Abstract
As robotic hardware and learning methods advance, humanoids need tools to perform tasks beyond their inherent physical limits. Successful tool use requires selecting a suitable tool and coordinating manipulation and, when needed, locomotion to complete the task. Existing benchmarks do not jointly evaluate these capabilities on a humanoid. We introduce HumanoidToolBench, an 18-task benchmark spanning three scenarios, three execution levels, and two tool-set modes, together with ToolBook, a dataset of 3.1k demonstrations collected in simulation and on a real Unitree G1. Evaluation of seven policies in simulation and three on the real robot reveals substantial gaps between selecting a suitable tool and completing the task. Focused GR00T N1.7 probes show reduced selection accuracy on unseen tools and continued task execution under unrelated instructions. Code and data are available at https://snu-pi.github.io/HumanoidToolBench/.