UniWAM: Unified World-Action Model

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of world dynamics grounding in vision-language-action models and the limited semantic reasoning of world-action models by proposing a unified architecture that integrates physical reasoning, world generation, and action prediction modules. Methodologically, low-level actions are represented via natural language, complemented by a complementary supervision pretraining strategy. Future visual noise augmentation and history-conditioned flow matching are introduced to reduce denoising steps. The approach further incorporates rigorous data curation, multi-source mixed training, and large-scale human-robot co-training. Experimental results demonstrate state-of-the-art performance across multidimensional evaluations, validate a log-linear scaling law for unified human-robot co-training, and yield significant improvements in robustness, generalization capability, and long-horizon task execution.
📝 Abstract
Vision-language-action models benefit from the understanding and reasoning capabilities of pretrained vision-language models, but action-only supervision provides limited grounding in world dynamics. Conversely, world-action models inherit spatiotemporal priors from video generation models, yet remain limited in semantic understanding and reasoning under distribution shifts. We introduce UniWAM, a unified architecture that integrates a physical reasoner, a world generator, and an action predictor to jointly learn semantic understanding of the physical world, visual generation, and action prediction. To ensure the quality of the training data, we developed a rigorous data cleaning and annotation pipeline for both human egocentric data and robot data. To adapt the vision-language component to embodied tasks while preserving its inherited language capabilities, we represent low-level actions in natural language and introduce a pre-training recipe that assigns complementary supervision from visual question answering (VQA) data, human egocentric data, and robot demonstrations to the appropriate model components. During post-training, future visual noise augmentation reduces reliance on precise future predictions, while history-conditioned flow matching uses encoded action history to initialize action generation. Together, these designs significantly reduce denoising steps while maintaining performance. UniWAM achieves state-of-the-art (SOTA) performance across multiple evaluations, including in-distribution performance, robustness, generalization, instruction following, and long-horizon task execution. Furthermore, we uncover a log-linear scaling law of unified human-robot co-training, demonstrating the effectiveness of large-scale pre-training on a mixture of human and robot data.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action Models
World-Action Models
Semantic Understanding
World Dynamics
Embodied AI
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified World-Action Model
History-Conditioned Flow Matching
Log-Linear Scaling Law
Vision-Language-Action
Embodied AI
🔎 Similar Papers
No similar papers found.