🤖 AI Summary
This study addresses the limitation of existing world action models in accurately predicting environmental dynamics due to the lack of structured, control-coupled world representations. To this end, we propose Magic-W0, which jointly models structured physical state evolution and continuous actions. Its core innovations include introducing the first structured world transition encompassing current states, 3D motion transitions, and future semantics; designing a hierarchically aligned interaction mechanism that enables bidirectional enhancement between prediction and control; and integrating visual-language context, 3D geometric supervision, and latent semantic modeling for large-scale pretraining. Experimental results demonstrate that Magic-W0 achieves state-of-the-art performance on the RoboDojo-Sim benchmark with a score of 27.10, while exhibiting strong generalization and rapid adaptation capabilities in real-world robotic tasks.
📝 Abstract
World-action models (WAMs) augment robot policies with action-conditioned environment dynamics, yet existing approaches largely rely on future observation reconstruction or generic latent prediction and lack structured, control-oriented world representations tightly coupled with action generation. We introduce Magic-W0, a world-action foundation model that jointly models structured physical state evolution and continuous actions. Magic-W0 represents interaction as a Structured World Transition consisting of Current State, Transition, and Future State. Current State combines vision-language context with Current 3D Geometry; Transition is represented by 3D Motion capturing action-induced three-dimensional changes; and Future State is represented by Future Semantics describing task-relevant outcomes. To couple prediction and control, we propose a layer-aligned world-action interaction architecture in which evolving action hypotheses condition world-transition prediction, while predicted world representations continuously inform action generation. Magic-W0 is pre-trained on large-scale egocentric human manipulation, UMI, real-robot, and simulation data, with latent supervision for geometry, 3D motion, and future semantics from pre-trained visual models. Inference-time interventions show that structured world representations respond systematically to changes in candidate actions and that action-related information propagates through shared 3D representations into future semantic predictions. On RoboDojo-Sim, Magic-W0 achieves an average Score of 27.10, the highest among the compared WAMs. Across multiple real-robot tasks, it also demonstrates strong downstream performance after fine-tuning with limited downstream data, supporting generalization and rapid adaptation.