Score
Designs, builds, and trains predictive models of environment dynamics that map past observations and agent actions to future latent states or observations, typically as action‑conditional or action‑observation world models often trained on video sequence data. This work specifies architectures and learning objectives (e.g., latent dynamics, reconstruction and prediction losses, attention‑driven retrieval, diffusion adaptations, or synthetic priors) and produces action‑conditioned rollouts and closed‑loop action sequences for use in downstream control or policy training.
Current world models suffer from conceptual ambiguity in embodied intelligence and generative simulation, lacking a unified classification and design framework tailored for robotic control. This work formally defines a world model as one conditioned on actions to predict the future evolution of task-relevant observations or states, and introduces a novel paradigm—world action models—that explicitly links prediction with executable actions. Building on this definition, the study systematically organizes four methodological families: imagine-then-execute, video-feature-conditioned action prediction, joint video-action modeling, and policy learning augmented by auxiliary video prediction. The research clarifies the conceptual boundaries of (action-conditioned) world models and presents the first structured taxonomy specifically designed for embodied prediction and control, thereby advancing standardized understanding and application in the field.
This study addresses the conceptual ambiguity surrounding World Action Models (WAMs) by clarifying their distinctions from related paradigms such as world models, video generation, and vision-language action policies. Through two complementary lenses—generated content type and methodological composition—the work proposes the first unified taxonomy to systematically deconstruct WAM design paradigms. It reveals a fundamental trade-off between representational richness and computational, memory, latency, and action annotation costs, highlighting a growing trend toward generating minimal yet control-critical future information. The paper characterizes WAMs as inherently predictive-action synergistic mechanisms, distills common design patterns, and provides a systematic overview of current advances and open challenges across key dimensions including interactivity, causality, persistence, physical plausibility, and generalization capability.
This work addresses the scarcity of action labels in world model training by proposing the Latent-Action World Model (LAWM), which jointly leverages a small set of action-labeled interaction trajectories and abundant unlabeled passive observations (e.g., videos). Methodologically, LAWM learns an action-invariant dynamical representation in a latent space, maps explicit actions into this latent action space, and performs cross-modal latent dynamics modeling via action-observation alignment. It innovatively unifies offline reinforcement learning with purely passive data training—marking the first approach to extend world model data sources and generalization capability under extreme action-label scarcity. Evaluated on the DeepMind Control Suite, LAWM achieves performance comparable to fully supervised baselines using only 10% of action annotations, demonstrating substantial improvements in data efficiency.
This work proposes A2World—the first action-conditioned diffusion-based world model—designed to learn transferable dynamics priors from large-scale, multi-view robotic manipulation data. By pretraining on action-driven visual scene evolution, A2World enables long-horizon, high-fidelity simulation and joint video-action prediction, and seamlessly integrates into specialized simulators and policy learning frameworks. Experimental results demonstrate that A2World significantly improves simulation fidelity and policy prediction accuracy, effectively substituting real-world trial-and-error with efficient what-if analysis. This approach establishes a new paradigm for enhancing the generalization capabilities of robotic systems through learned world models.
Existing world models rely heavily on large-scale labeled action data and computationally expensive training, hindering rapid adaptation to novel environments with heterogeneous action spaces and scarce annotations. To address this, we propose a self-supervised framework that eliminates the need for explicit action labels: first, video representation learning implicitly extracts action representations from inter-frame dynamics; second, an autoregressive world model is constructed conditioned on these latent actions. This constitutes the first approach to integrate action modeling directly into the world model pretraining stage, enabling action-agnostic universal representation learning. Our method achieves cross-action-space transfer with only minimal environment interaction. Extensive experiments across multiple environments demonstrate substantial improvements in video prediction fidelity and visual planning performance, reduce fine-tuning costs by over 40%, and exhibit strong generalization across diverse action spaces.
This work proposes LingBot-VA, an autoregressive diffusion-based control framework that integrates video world modeling with causal reasoning to enhance long-horizon robotic control and generalization in complex environments. By leveraging a Mixture-of-Transformers architecture, the method constructs a shared latent space for vision and action, enabling joint learning of video frame prediction and policy execution. It further incorporates closed-loop rolling inference and asynchronous parallel control mechanisms to improve temporal coherence and responsiveness. Experimental results demonstrate that LingBot-VA significantly outperforms baseline approaches in both simulation and real-world settings, achieving higher success rates on long-horizon tasks, improved data efficiency, and stronger generalization to novel environmental configurations.
This work investigates the identifiability of latent states and dynamics in controlled world models under high-dimensional observations and restricted behavioral policies. By introducing two key conditions—spectral separability and non-degenerate action variation—it establishes, for the first time, that the global optimum of Joint Embedding Predictive Architecture (JEPA) recovers the true latent states and dynamics uniquely up to an orthogonal transformation. The paper further provides quantitative identifiability bounds under approximate optimization. The theoretical analysis integrates Gaussian latent modeling, spectral methods, and counterfactual perturbation constructions. Empirical results demonstrate that insufficient action coverage adversely impacts transition identifiability and counterfactual prediction accuracy, while the proposed approach significantly enhances goal-directed planning performance in latent space.
This work addresses the challenge of unifying prediction, planning, and irreversibility within world models. It formulates prediction as a probability measure over future trajectories and, under a local Markov assumption, employs the Onsager–Machlup action functional to decompose latent dynamics into reversible and irreversible components in path space. The authors introduce rollout-based entropy production as an operational measure of irreversibility. Through path-integral analysis, attention mechanism inspection, and small-scale model experiments, they find that attention asymmetry emerges in response to increasing data irreversibility. While symmetrization interventions suppress entropy production, they selectively impair long-horizon prediction of irreversible processes yet preserve the model’s capacity to capture relaxation dynamics.
Existing world models for robotic action struggle to simultaneously achieve efficient action prediction and explicit dynamics modeling during inference, often hindered by high computational costs or the absence of future state representations. This work proposes ForeWAM, an implicit future-state-driven direct-policy world model that provides dynamic context for action decisions without generating future videos. Its key innovations include a Future-KV mechanism that reuses key-value states from both visual inputs and future slots, and a dynamics register supervised by a frozen implicit action teacher to implicitly capture interaction-induced dynamics. Integrating Video DiT pre-filling with cross-layer state reuse, ForeWAM achieves 96.7% (standard) and 96.9% (accelerated) success rates on the LIBERO benchmark and 61.6% on LIBERO-Plus, all without requiring embodied pretraining.
This work addresses a critical limitation in existing latent-variable world models, which rely on average prediction error over training data for training and selection—a metric that fails to reflect actual controller performance due to a mismatch between the evaluation distribution and the distribution queried by the planner. The authors propose instead to center model assessment on the discrepancy between predicted and true costs over states reachable by the planner. They establish, for the first time, a rigorous theoretical link between this discrepancy and control suboptimality, proving it provides a valid upper bound on performance loss, whereas conventional prediction errors neither bound nor track performance. Leveraging control theory, spectral analysis, and non-normal operator theory, they decompose the discrepancy into an intrinsic manifold residual and an off-manifold divergence term, and introduce a fidelity score to quantify alignment of the planner’s reachable distribution. Experiments on synthetic systems and model predictive control confirm that the proposed metric reliably tracks control performance, while single-step prediction error shows virtually no correlation.