OpenWAM: An Open Framework for Composable World-Action Models

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the difficulty of comparing architectural differences in existing world action models due to variable coupling. We propose an open-source framework built upon a causal video foundation, which unifies joint generation, sequential generation, and decoupled inference through configurable video-action interaction structures. Methodologically, we design a shared hybrid Transformer to enable modular interactions and introduce counterfactual supervision to optimize dynamics learning for improved generalization. Experimental results demonstrate that the proposed framework achieves a 97.8% success rate on the LIBERO benchmark. Furthermore, when adapting only the video predictor while keeping the inverse dynamics model frozen, it attains 84.0% performance on unseen tasks, significantly outperforming existing baselines.
📝 Abstract
World-action models (WAMs) couple future prediction with robot control, yet existing systems often vary the video backbone, interaction structure, supervision, and inference procedure simultaneously, making their design choices difficult to compare. We introduce OPENWAM, an open world-action modeling framework built around a common causal robot-video foundation and configurable video-action interaction. Starting from Wan2.2-5B, we perform causal robot-video pretraining on over 10,000 hours of video, then integrate an action expert through a shared Mixture-of-Transformers architecture that supports joint, video-then-action, action-then-video, and decoupled generation. OPENWAM achieves high success rates on four LIBERO suites and real-world bimanual tasks; robot-video training with causal adaptation improves VTA success on LIBERO-Long from 68.4% to 97.8%. The same configurable architecture naturally extends to inverse and forward dynamics, allowing us to study how counterfactual transitions improve independently trained dynamics models beyond demonstrations alone. When only the video predictor is adapted to a new task, a frozen local-context inverse dynamics model trained on counterfactual data and demonstrations achieves 84.0% mean success across four held-out LIBERO-90 tasks, compared with 47.0% for a full-context inverse model and 21.5% for a local-context model trained only on demonstrations. For forward dynamics, counterfactual supervision reduces RGB prediction error by 34.5% and raises outcome identification from 21.1% to 71.3% among 16 same-state outcomes. OPENWAM provides a common testbed for comparing WAM interaction designs and for studying dynamics learning from video data beyond successful demonstrations.
Problem

Research questions and friction points this paper is trying to address.

World-Action Models
comparability
dynamics learning
robot control
Innovation

Methods, ideas, or system contributions that make the work stand out.

World-Action Models
Mixture-of-Transformers
Causal Robot-Video Pretraining
Counterfactual Transitions
Inverse and Forward Dynamics
🔎 Similar Papers
No similar papers found.