World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing vision-language models (VLMs) in lacking an embodied action perspective for robotic manipulation by proposing a multi-agent framework. The method constructs a visual-action workspace that enables direct manipulation through geometric contact views, world-action rehearsal, and closed-loop correction. Furthermore, it establishes an embodied knowledge acquisition paradigm that evolves skills from expert videos and trains compact VLMs using interaction trajectories. Experimental results demonstrate that the proposed framework achieves a 75.6% success rate on the LIBERO-Pro benchmark, significantly outperforming baselines. Notably, after fine-tuning, the Qwen3.5-9B model exhibits a substantial improvement in out-of-distribution task success rate, rising from 1.7% to 43.3%, thereby validating both the effectiveness and generalization capability of the proposed approach.
📝 Abstract
General-purpose vision-language models (VLMs) bring broad knowledge and spatial reasoning to robot manipulation, yet existing systems either use them indirectly, to predict constraints or write programs, or give them a view of the scene rather than a world in which to act. We present World Action Agent (WAA), a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace. The workspace has three properties. Contact views, selected automatically from the scene geometry, present the scene around the current interaction. Action rehearsal turns each action into an editable proposal that the agent, alone or through an Imagination Agent, previews and revises against planning feedback before execution. In-view correction closes the loop between observation, rehearsal, and low-level execution, letting the agent remove residual offsets in the view where it observes them. Through the same workspace, WAA acquires embodied procedural knowledge in two ways: it evolves multimodal skills from expert videos and human teaching under evidence-based review and consults them through a Skill Agent, and its interaction traces train smaller VLMs to pilot the same harness. On LIBERO-Pro, WAA with skills evolved only from LIBERO-90 reaches a state-of-the-art 75.6% average success, outperforming end-to-end VLAs, code-as-policy agents, and a visual-harness baseline with the same backbone; the same skills remain effective on robosuite without further learning. Fine-tuning Qwen3.5-9B on harness traces raises its out-of-domain success from 1.7% to 43.3%.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Robot Manipulation
Embodied AI
Action Rehearsal
Skill Acquisition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Action Rehearsal
Multi-Agent System
Robot Manipulation
Embodied Procedural Knowledge
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yehang Zhang
HKUST(GZ)
H
Haojian Huang
HKUST(GZ)
Y
Yifan Chang
Knowin AI
J
Jianchong Su
HKUST(GZ)
B
Bohan Zhou
CUHK
Yingjie Xu
Yingjie Xu
Hong Kong University of Science and Technology(Guang Zhou))
Computer Vision
W
Wosong Chen
HKUST(GZ)
T
Tianhao Zhou
HKUST(GZ)
C
Chenxu Wang
Knowin AI
T
Tianyi Zhang
Knowin AI
Y
Yangkai Wei
Knowin AI
W
Wenqian Li
CUHK
S
Shiyuan Deng
Knowin AI
Yinchuan Li
Yinchuan Li
Principal Researcher, Noah's Ark Lab
Generative ModelsEmbodied AIArtificial Intelligence
Ying-Cong Chen
Ying-Cong Chen
Hong Kong University of Science and Technology (Guangzhou)
Computer Vision and Pattern Recognition
Zexi Li
Zexi Li
Alibaba Group
Deep LearningLarge Language ModelsFederated Learning