MiniWAM: Learning Compact Future Targets for Efficient World-Action Modeling

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive training costs and inefficiency of existing world-action models caused by predicting high-dimensional native visual features. To this end, we propose PRISM, a method that learns compact future predictive representations as surrogate targets to enable efficient joint world-action modeling. By integrating inverse dynamics supervision with feature reconstruction mechanisms, PRISM significantly compresses future state dimensionality while preserving essential control information. Furthermore, it constructs a compact representation space based on DINOv3 and WAN2.1 VAE through inverse spatiotemporal modeling, frozen encoders, and joint prediction strategies. Experimental results demonstrate that PRISM accelerates training by 8× and reduces feature tokens by 65×, outperforming larger models across multiple simulation benchmarks despite utilizing fewer parameters.
📝 Abstract
World modeling has emerged as an effective co-training objective for robot policies, giving rise to World Action Models (WAMs) that jointly predict actions and future states. However, most WAMs predict future states in the native representation space of pretrained visual backbones, resulting in high-dimensional targets with substantial training cost. We introduce MiniWAM, which instead predicts compact future representations learned from privileged current-future transitions. To construct these targets, we propose Predictive Representations via Inverse Spatiotemporal Modeling (PRISM), which combines inverse-dynamics supervision with feature reconstruction to emphasize control-relevant transition information while preserving useful future-state information. With the learned PRISM encoder frozen, MiniWAM is trained to jointly predict the resulting targets and robot actions from current observations. With 65$\times$ fewer native future feature tokens, MiniWAM consistently outperforms native future-feature prediction with both DINOv3 and WAN2.1 VAE features, while achieving up to an 8$\times$ speedup in world-action training. At 0.25B parameters, MiniWAM is already competitive with substantially larger WAMs on LIBERO, LIBERO-Plus, and RoboTwin 2.0 simulation benchmarks. Representation analyses further show that PRISM contributes behavioral structure beyond feature reconstruction alone. These results demonstrate that effective world-action modeling does not require predicting native visual futures, and that compact predictive representations provide a strong and substantially more efficient target for policy learning. The project page is available at: https://j1dan.github.io/MiniWAM.
Problem

Research questions and friction points this paper is trying to address.

World-Action Modeling
High-dimensional targets
Training efficiency
Future state prediction
Robot policies
Innovation

Methods, ideas, or system contributions that make the work stand out.

World-Action Modeling
Compact Predictive Representations
Inverse Spatiotemporal Modeling
Robot Policy Learning
Efficient Training
🔎 Similar Papers
J
Jie Chen
Department of Mechanical Engineering, National University of Singapore, Singapore
R
Ruofei Bai
Nanyang Technological University, Singapore
Y
Yuxin Cai
Nanyang Technological University, Singapore
Y
Yifeng Zhang
Department of Mechanical Engineering, National University of Singapore, Singapore
C
Chengyang He
Department of Mechanical Engineering, National University of Singapore, Singapore
J
Jun Li
A*STAR Institute of Advanced Intelligence and Computing (A*STAR IAIC), Singapore
W
Wei-Yun Yau
A*STAR Institute of Advanced Intelligence and Computing (A*STAR IAIC), Singapore
Guillaume Sartoretti
Guillaume Sartoretti
Assistant Professor, National University of Singapore (NUS), Mechanical Engineering Dpt
Multi-Agent SystemsRoboticsSwarm IntelligenceDistributed ControlDistributed Learning