🤖 AI Summary
Current vision-language models (VLMs) exhibit poor performance in embodied tasks, yet it remains challenging to disentangle whether failures stem from flawed high-level decision-making or inadequate low-level motor control. To address this, this work proposes HumanCLAW, a novel framework that decouples high-level action planning from physical execution: the VLM outputs only atomic skill commands, which are then mapped by the system into physically plausible full-body motions. Leveraging this paradigm, we introduce HumanCLAW-Bench, a benchmark comprising 1,218 long-horizon embodied tasks. Experiments across nine state-of-the-art VLMs reveal a maximum success rate of merely 16.8%, exposing a critical deficiency—these models lack sustained awareness of their own bodily states. This finding underscores embodied self-awareness as a fundamental bottleneck in current VLMs for real-world interaction.
📝 Abstract
Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed to execute it, e.g., losing balance and falling. In this work, we introduce HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution. At every step, a harnessed, off-the-shelf VLM issues an atomic skill command, and the command is translated into a sub-second chunk of continuous full-body motion with real physical consequences, including gravity and collisions. The body can therefore act freely in the physical world, while execution-side disturbances, balance and motor errors, are factored out. What remains measurable is the model's action intelligence: its moment-to-moment choice of what the body should execute next. Based on this framework, we build HumanCLAW-Bench: 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes. We test nine state-of-the-art VLMs and find that none solves the benchmark; the best model reaches only a 16.8% success rate. Recognizing the target is not the bottleneck. What current VLMs lack is embodied self-awareness: they lose track of their own body, failing to tell where it is, whether it has reached the goal, or whether it has hit an obstacle.