From World Models to World Action Models: Rethinking Next-State Prediction

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing world action models (WAMs) that rely on fixed next-state representations, which constrains the inductive bias for action learning and hinders cross-embodiment generalization. To overcome this, we propose CF-WAM, a dynamic prediction framework that samples multi-dimensional projections—including visual and semantic representations—and normalizes them into a unified video format. This establishes a common reference frame to eliminate appearance discrepancies across embodiments. By integrating a joint multi-projection supervision mechanism, the framework accurately captures state transition structures, enabling unified WAM training. Experimental results demonstrate that CF-WAM substantially improves both training efficiency and control performance, achieving success rates of 82.5% and 82.65% on the RoboCasa and LIBERO benchmarks, respectively, thereby effectively realizing robust cross-embodiment generalization.
📝 Abstract
Predicting the next state is a core paradigm of World Models for modeling physical dynamics, emphasizing prediction fidelity. As World Models evolve into World-Action Models (WAMs), existing methods still fix the next state before training as RGB, a single latent feature, or a static combination of predefined targets, thereby constraining action learning to the inductive biases preserved by a particular representation. To address this limitation, we propose CF-WAM, a dynamic next-state prediction framework that samples visual, semantic, geometric, and interaction projections of the same future, standardizes them into a common video form, and supervises a unified WAM across these projections. The action-relevant constraints exposed by these projections accumulate across training steps, forcing WAM to capture the underlying state-transition structure that supports multiple projections of the same action-conditioned future. This dynamic mechanism also provides a natural cross-embodiment dynamics reference frame for Human and Robot learning. By jointly learning across different next-state parameterizations, heterogeneous Human and Robot experience can bypass appearance differences and directly contribute to shared state-transition learning, improving cross-embodiment generalization. Experiments show that CF-WAM improves both training efficiency and final control performance, while translating Human experience effectively into policy gains. CF-WAM achieves state-of-the-art performance on RoboCasa-GR1 with an average success rate of 82.50%, while reaching 82.65% on LIBERO-Plus and up to 84.00% in real-world evaluations.
Problem

Research questions and friction points this paper is trying to address.

World-Action Models
Next-State Prediction
Cross-Embodiment Generalization
Inductive Bias
State-Transition Structure
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Action Models
Dynamic Next-State Prediction
Cross-Embodiment Learning
Multi-Projection Supervision
State-Transition Structure
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
T
Tingyu Yuan
CASIA
Z
Ziming Ji
BUPT
B
Biaoliang Guan
XJTU
W
Wen Ye
CASIA
W
Wenrui Tian
WHU
Z
Zhaopeng Gu
CASIA
F
Feihong Zhang
THU
X
Xu Yang
CASIA
Y
Yan Huang
CASIA
Zhaowen Li
Zhaowen Li
National Laboratory of Pattern Recognition,Institute of Automation,Chinese Academy of Sciences
Computer VisionArtificial IntelligenceSelf-supervised Learning
Chaoyang Zhao
Chaoyang Zhao
Institute of Automation, Chinese Academy of Sciences
computer vision
J
Jinqiao Wang
CASIA