HumanCLAW: Can Vision-Language Models Act Through a Body?

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current vision-language models (VLMs) exhibit poor performance in embodied tasks, yet it remains challenging to disentangle whether failures stem from flawed high-level decision-making or inadequate low-level motor control. To address this, this work proposes HumanCLAW, a novel framework that decouples high-level action planning from physical execution: the VLM outputs only atomic skill commands, which are then mapped by the system into physically plausible full-body motions. Leveraging this paradigm, we introduce HumanCLAW-Bench, a benchmark comprising 1,218 long-horizon embodied tasks. Experiments across nine state-of-the-art VLMs reveal a maximum success rate of merely 16.8%, exposing a critical deficiency—these models lack sustained awareness of their own bodily states. This finding underscores embodied self-awareness as a fundamental bottleneck in current VLMs for real-world interaction.
📝 Abstract
Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed to execute it, e.g., losing balance and falling. In this work, we introduce HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution. At every step, a harnessed, off-the-shelf VLM issues an atomic skill command, and the command is translated into a sub-second chunk of continuous full-body motion with real physical consequences, including gravity and collisions. The body can therefore act freely in the physical world, while execution-side disturbances, balance and motor errors, are factored out. What remains measurable is the model's action intelligence: its moment-to-moment choice of what the body should execute next. Based on this framework, we build HumanCLAW-Bench: 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes. We test nine state-of-the-art VLMs and find that none solves the benchmark; the best model reaches only a 16.8% success rate. Recognizing the target is not the bottleneck. What current VLMs lack is embodied self-awareness: they lose track of their own body, failing to tell where it is, whether it has reached the goal, or whether it has hit an obstacle.
Problem

Research questions and friction points this paper is trying to address.

vision-language models
embodied action
physical embodiment
action intelligence
embodied self-awareness
Innovation

Methods, ideas, or system contributions that make the work stand out.

embodied AI
vision-language models
action decision-making
physical simulation
self-awareness
S
Siyao Li
Meta
Jiawei Gu
Jiawei Gu
Sun Yat-sen University
Natural language processingMultimodal reasoning
Shuai Liu
Shuai Liu
School of Electrical and Electronic Engineering, Nanyang Technological University
OptimizationOptimal controlTime Delay SystemMulti-agent SystemSignal Processing
K
Kairui Hu
Nanyang Technological University
Z
Zekun Li
Meta
Linjie Li
Linjie Li
Microsoft
Vision and Language
Chengcheng Tang
Chengcheng Tang
Meta
Computer GraphicsGeometric Computing
Po-Chen Wu
Po-Chen Wu
Research Scientist at Reality Labs
Computer VisionAugmented RealityRoboticsHuman–computer Interaction
Ivan Shugurov
Ivan Shugurov
Technische Universität München
machine learningcomputer vision
Lingni Ma
Lingni Ma
Meta Reality Labs Research
computer visionmachine learningtracking
Michael Zollhoefer
Michael Zollhoefer
Director, Research Scientist, Reality Labs Research, Meta
Neural RenderingComputer VisionMachine LearningComputer Graphics
S
Sizhe An
Meta
A
Abhay Mittal
Meta
Amy Zhao
Amy Zhao
Research Scientist, Facebook Reality Labs
computer visioncomputer graphicsmachine learning
Ranjay Krishna
Ranjay Krishna
University of Washington, Allen Institute for AI
Computer VisionNatural Language ProcessingMachine LearningHuman Computer Interaction
Manling Li
Manling Li
Assistant Professor at Northwestern University
Natural Language ProcessingVision-LanguageEmbodied Agents
Ziwei Liu
Ziwei Liu
Associate Professor, Nanyang Technological University
Computer VisionMachine LearningComputer Graphics
Chuan Guo
Chuan Guo
Research Scientist, Reality Labs Research @ Meta
3D AnimationHuman Motion SynthesisDeep Generative Model