Streaming-WAM: Action-Conditioned World-Action Model for Asynchronous Robot Manipulation

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high inference costs of world action models and the failure of visual predictions to account for committed actions during asynchronous execution. To overcome these limitations, we propose Streaming-WAM, a framework that integrates action-conditioned world modeling with asynchronous control. This approach introduces a streaming update mechanism that enables visual predictions to perceive previously committed actions, while leveraging a fixed prefix to guide subsequent action generation for efficient and precise manipulation under asynchronous execution. Experimental results demonstrate that Streaming-WAM achieves a 98.35% success rate on the LIBERO benchmark while reducing inference latency by 2.93×. Furthermore, in real-world robotic tasks, the total execution time is decreased from 90 seconds to 38 seconds, validating the effectiveness of the proposed method in both simulated and physical environments.
📝 Abstract
World action models (WAMs) that use future visual prediction at inference time incur substantial generation costs. Asynchronous execution reduces waiting by overlapping inference with robot motion, but visual predictions used for subsequent action generation must anticipate the effects of actions already scheduled for execution during inference. We introduce Streaming-WAM, which couples action-conditioned world modeling with asynchronous robot control to account for committed actions in future visual prediction. At each streaming update, the model conditions future visual prediction on the latest observation and the committed actions, which form the fixed prefix of the next action chunk. The resulting action-conditioned visual features guide generation of the remaining actions within the same joint update, so the continuation is informed by the scene changes expected during execution of the fixed prefix. On LIBERO, Streaming-WAM achieves an average success rate of 98.35\% and reduces mean episode time by a factor of 2.93 relative to Fast-WAM. On the real-world Stamp Paper task, mean episode time falls from 90 s with synchronous Joint-WAM to 38 s with Streaming-WAM. These results show that Streaming-WAM supports efficient asynchronous control while maintaining high task success rates.
Problem

Research questions and friction points this paper is trying to address.

World Action Models
Asynchronous Robot Manipulation
Visual Prediction
Action-Conditioned Modeling
Inference Latency
Innovation

Methods, ideas, or system contributions that make the work stand out.

World-Action Model
Asynchronous Robot Control
Action-Conditioned Visual Prediction
Streaming Update
Robot Manipulation
X
Xuyao Huang
SJTU
Y
Yixuan Wang
SJTU
Z
Zengyao Ye
SDU
B
Boyuan Zhao
ECUST
Chenyang Yu
Chenyang Yu
Dalian University of Technology
Deep learning,person reidentification
H
Haoran Wen
Li Auto Inc.
Z
Zhijie Deng
SJTU