Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing world action models in handling long-horizon compositional manipulation due to their lack of explicit subtask reasoning. To this end, this work proposes ViGAR, a hierarchical framework that decomposes manipulation into a visual subgoal planner and an executor, enabling joint trajectory and action generation through shared pretrained world model representations. The framework introduces a hierarchical subgoal reasoning mechanism that supports in-context learning from global goal images, allowing dynamic adjustment of behavior decomposition without parameter updates. Evaluated on the RoboTwin benchmark, ViGAR achieves an average success rate improvement of 12.86% over baselines. Real-world robot experiments further validate its effectiveness in both compositional and in-context learning tasks.
📝 Abstract
Long-horizon compositional manipulation has become increasingly important for real-world robot deployment, where a single task involves multiple coordinated subtasks. Existing world-action models (WAMs) jointly predict short-horizon visual futures and actions, but typically lack explicit subtask-level reasoning. We propose Visual Goal-conditioned Action Reasoning (ViGAR), a hierarchical framework that factorizes manipulation into a visual subgoal planner and a subgoal executor. Given the current observation and global instruction, the subgoal planner predicts a visual subgoal for the next subtask. The subgoal executor then jointly generates future visual trajectories and actions conditioned on the predicted subgoal. Both components share a pretrained world-model representation, enabling task-level planning and action generation to benefit from common physical knowledge. Moreover, our framework naturally supports in-context learning: using a global goal image as context can induce different subtask decompositions and behaviors without parameter updates. On the RoboTwin Clean2Random benchmark, ViGAR achieves 82.00% and 67.02% success rates under the Clean and Random settings, respectively, surpassing the strongest baseline by 12.86 percentage points in average success rate. Real-world robot experiments on five compositional and two in-context learning tasks further confirm the effectiveness of ViGAR.
Problem

Research questions and friction points this paper is trying to address.

Long-horizon compositional manipulation
World-action models
Subtask-level reasoning
In-context learning
Robotic manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

World-Action Model
Hierarchical Framework
Visual Subgoal Planning
In-Context Learning
Compositional Manipulation
S
Shukai Gong
Peking University
X
Xuanran Zhai
AgiBot
Y
Yintianrun Zhang
Peking University
R
Ruopeng Cui
AgiBot
Y
Ye Huang
Peking University
Y
Yiyang Fu
Peking University
D
Dexuan Lyu
AgiBot
C
Chaojie Li
AgiBot
X
Xinyi Song
AgiBot
P
Peiwen Lin
AgiBot
C
Chuang Wang
AgiBot
M
Mingyuan Jia
CocoMatrix
Yufan Deng
Yufan Deng
Oxford VGG
J
Jiaxin Fang
CocoMatrix
B
Bo Liang
Peking University
J
Jiaxin Li
Peking University
Yuxiang Gao
Yuxiang Gao
Johns Hopkins University
RoboticsHuman-Robot InteractionSocially-aware Navigation
H
Hao Liu
AgiBot
Daquan Zhou
Daquan Zhou
Bytedance, US
Artificial IntelligenceDeep learning