🤖 AI Summary
This study addresses the prohibitive training costs and inefficiency of existing world-action models caused by predicting high-dimensional native visual features. To this end, we propose PRISM, a method that learns compact future predictive representations as surrogate targets to enable efficient joint world-action modeling. By integrating inverse dynamics supervision with feature reconstruction mechanisms, PRISM significantly compresses future state dimensionality while preserving essential control information. Furthermore, it constructs a compact representation space based on DINOv3 and WAN2.1 VAE through inverse spatiotemporal modeling, frozen encoders, and joint prediction strategies. Experimental results demonstrate that PRISM accelerates training by 8× and reduces feature tokens by 65×, outperforming larger models across multiple simulation benchmarks despite utilizing fewer parameters.
📝 Abstract
World modeling has emerged as an effective co-training objective for robot policies, giving rise to World Action Models (WAMs) that jointly predict actions and future states. However, most WAMs predict future states in the native representation space of pretrained visual backbones, resulting in high-dimensional targets with substantial training cost. We introduce MiniWAM, which instead predicts compact future representations learned from privileged current-future transitions. To construct these targets, we propose Predictive Representations via Inverse Spatiotemporal Modeling (PRISM), which combines inverse-dynamics supervision with feature reconstruction to emphasize control-relevant transition information while preserving useful future-state information. With the learned PRISM encoder frozen, MiniWAM is trained to jointly predict the resulting targets and robot actions from current observations. With 65$\times$ fewer native future feature tokens, MiniWAM consistently outperforms native future-feature prediction with both DINOv3 and WAN2.1 VAE features, while achieving up to an 8$\times$ speedup in world-action training. At 0.25B parameters, MiniWAM is already competitive with substantially larger WAMs on LIBERO, LIBERO-Plus, and RoboTwin 2.0 simulation benchmarks. Representation analyses further show that PRISM contributes behavioral structure beyond feature reconstruction alone. These results demonstrate that effective world-action modeling does not require predicting native visual futures, and that compact predictive representations provide a strong and substantially more efficient target for policy learning. The project page is available at: https://j1dan.github.io/MiniWAM.