🤖 AI Summary
This study addresses the lack of first-person visual evaluation in existing multi-agent collaboration benchmarks by introducing a novel simulation benchmark tailored for first-person perspectives. Methodologically, we construct ten tasks encompassing diverse collaborative modes and collect synchronized demonstration data through a multi-operator VR teleoperation pipeline to facilitate policy development and evaluation. Experiments assessing representative visuomotor policies reveal core challenges in coordinating perception and control. By bridging this critical evaluation gap, this work establishes a robust data and assessment foundation for advancing research on multi-agent collaborative strategies.
📝 Abstract
Many physical tasks in human environments require collaboration, from assisting a partner to jointly manipulating an object. Yet, existing humanoid benchmarks largely focus on single-humanoid skills and lack evaluation of multi-humanoid collaboration under egocentric visual observations. We introduce CoHuB (Collaborative Multi-Humanoid Benchmark), a simulation benchmark for multi-humanoid collaboration under egocentric visual observations. CoHuB provides 10 tasks, eight with two humanoids and two with three humanoids, spanning diverse collaboration patterns. We also provide synchronized demonstrations collected through a multi-operator VR teleoperation pipeline, in which each operator controls one humanoid from its egocentric view. Experiments with representative visuomotor policies reveal substantial challenges across different forms of coordinated perception and control. CoHuB provides a foundation for developing and evaluating multi-humanoid collaboration policies.