train world models

Designs, builds, and trains predictive models of environment dynamics that map past observations and agent actions to future latent states or observations, typically as action‑conditional or action‑observation world models often trained on video sequence data. This work specifies architectures and learning objectives (e.g., latent dynamics, reconstruction and prediction losses, attention‑driven retrieval, diffusion adaptations, or synthetic priors) and produces action‑conditioned rollouts and closed‑loop action sequences for use in downstream control or policy training.

trainworldmodels

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.48
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

This study addresses the conceptual ambiguity surrounding World Action Models (WAMs) by clarifying their distinctions from related paradigms such as world models, video generation, and vision-language action policies. Through two complementary lenses—generated content type and methodological composition—the work proposes the first unified taxonomy to systematically deconstruct WAM design paradigms. It reveals a fundamental trade-off between representational richness and computational, memory, latency, and action annotation costs, highlighting a growing trend toward generating minimal yet control-critical future information. The paper characterizes WAMs as inherently predictive-action synergistic mechanisms, distills common design patterns, and provides a systematic overview of current advances and open challenges across key dimensions including interactivity, causality, persistence, physical plausibility, and generalization capability.

action groundingembodied AIvideo generation

Must-Read Papers

Most classic and influential ideas
View more

Latent Action World Models for Control with Unlabeled Trajectories

Dec 10, 2025
MA
Marvin Alles
🏛️ Technical University of Munich | Eötvös Loránd University Budapest | Foundation Robotics Labs | Machine Learning Research Lab | Volkswagen Group

This work addresses the scarcity of action labels in world model training by proposing the Latent-Action World Model (LAWM), which jointly leverages a small set of action-labeled interaction trajectories and abundant unlabeled passive observations (e.g., videos). Methodologically, LAWM learns an action-invariant dynamical representation in a latent space, maps explicit actions into this latent action space, and performs cross-modal latent dynamics modeling via action-observation alignment. It innovatively unifies offline reinforcement learning with purely passive data training—marking the first approach to extend world model data sources and generalization capability under extreme action-label scarcity. Evaluated on the DeepMind Control Suite, LAWM achieves performance comparable to fully supervised baselines using only 10% of action annotations, demonstrating substantial improvements in data efficiency.

Enables offline RL using large-scale unlabeled and few labeled samples.Learns world models from both action-labeled and action-free data.Reduces reliance on scarce action-labeled trajectories for training.

This work proposes A2World—the first action-conditioned diffusion-based world model—designed to learn transferable dynamics priors from large-scale, multi-view robotic manipulation data. By pretraining on action-driven visual scene evolution, A2World enables long-horizon, high-fidelity simulation and joint video-action prediction, and seamlessly integrates into specialized simulators and policy learning frameworks. Experimental results demonstrate that A2World significantly improves simulation fidelity and policy prediction accuracy, effectively substituting real-world trial-and-error with efficient what-if analysis. This approach establishes a new paradigm for enhancing the generalization capabilities of robotic systems through learned world models.

action-conditioneddynamics priorsrobot learning

AdaWorld: Learning Adaptable World Models with Latent Actions

Mar 24, 2025
SG
Shenyuan Gao
🏛️ HKUST | Harvard University | Google DeepMind | UMass Amherst | MIT-IBM Watson AI Lab

Existing world models rely heavily on large-scale labeled action data and computationally expensive training, hindering rapid adaptation to novel environments with heterogeneous action spaces and scarce annotations. To address this, we propose a self-supervised framework that eliminates the need for explicit action labels: first, video representation learning implicitly extracts action representations from inter-frame dynamics; second, an autoregressive world model is constructed conditioned on these latent actions. This constitutes the first approach to integrate action modeling directly into the world model pretraining stage, enabling action-agnostic universal representation learning. Our method achieves cross-action-space transfer with only minimal environment interaction. Extensive experiments across multiple environments demonstrate substantial improvements in video prediction fidelity and visual planning performance, reduce fine-tuning costs by over 40%, and exhibit strong generalization across diverse action spaces.

Enabling efficient adaptation to novel environments with limited interactionsLearning adaptable world models with latent actionsReducing reliance on action-labeled data and costly training

This work proposes LingBot-VA, an autoregressive diffusion-based control framework that integrates video world modeling with causal reasoning to enhance long-horizon robotic control and generalization in complex environments. By leveraging a Mixture-of-Transformers architecture, the method constructs a shared latent space for vision and action, enabling joint learning of video frame prediction and policy execution. It further incorporates closed-loop rolling inference and asynchronous parallel control mechanisms to improve temporal coherence and responsiveness. Experimental results demonstrate that LingBot-VA significantly outperforms baseline approaches in both simulation and real-world settings, achieving higher success rates on long-horizon tasks, improved data efficiency, and stronger generalization to novel environmental configurations.

action-visual dynamicscausal world modelinglong-horizon manipulation

Latest Papers

What's happening recently
View more

This work investigates the identifiability of latent states and dynamics in controlled world models under high-dimensional observations and restricted behavioral policies. By introducing two key conditions—spectral separability and non-degenerate action variation—it establishes, for the first time, that the global optimum of Joint Embedding Predictive Architecture (JEPA) recovers the true latent states and dynamics uniquely up to an orthogonal transformation. The paper further provides quantitative identifiability bounds under approximate optimization. The theoretical analysis integrates Gaussian latent modeling, spectral methods, and counterfactual perturbation constructions. Empirical results demonstrate that insufficient action coverage adversely impacts transition identifiability and counterfactual prediction accuracy, while the proposed approach significantly enhances goal-directed planning performance in latent space.

action-conditioned predictionbehavior policycontrolled world models

This work addresses the challenge of unifying prediction, planning, and irreversibility within world models. It formulates prediction as a probability measure over future trajectories and, under a local Markov assumption, employs the Onsager–Machlup action functional to decompose latent dynamics into reversible and irreversible components in path space. The authors introduce rollout-based entropy production as an operational measure of irreversibility. Through path-integral analysis, attention mechanism inspection, and small-scale model experiments, they find that attention asymmetry emerges in response to increasing data irreversibility. While symmetrization interventions suppress entropy production, they selectively impair long-horizon prediction of irreversible processes yet preserve the model’s capacity to capture relaxation dynamics.

entropy productionirreversibilitypath-space

Existing world models for robotic action struggle to simultaneously achieve efficient action prediction and explicit dynamics modeling during inference, often hindered by high computational costs or the absence of future state representations. This work proposes ForeWAM, an implicit future-state-driven direct-policy world model that provides dynamic context for action decisions without generating future videos. Its key innovations include a Future-KV mechanism that reuses key-value states from both visual inputs and future slots, and a dynamics register supervised by a frozen implicit action teacher to implicitly capture interaction-induced dynamics. Integrating Video DiT pre-filling with cross-layer state reuse, ForeWAM achieves 96.7% (standard) and 96.9% (accelerated) success rates on the LIBERO benchmark and 61.6% on LIBERO-Plus, all without requiring embodied pretraining.

action generationfuture video predictioninference efficiency

This work addresses a critical limitation in existing latent-variable world models, which rely on average prediction error over training data for training and selection—a metric that fails to reflect actual controller performance due to a mismatch between the evaluation distribution and the distribution queried by the planner. The authors propose instead to center model assessment on the discrepancy between predicted and true costs over states reachable by the planner. They establish, for the first time, a rigorous theoretical link between this discrepancy and control suboptimality, proving it provides a valid upper bound on performance loss, whereas conventional prediction errors neither bound nor track performance. Leveraging control theory, spectral analysis, and non-normal operator theory, they decompose the discrepancy into an intrinsic manifold residual and an off-manifold divergence term, and introduce a fidelity score to quantify alignment of the planner’s reachable distribution. Experiments on synthetic systems and model predictive control confirm that the proposed metric reliably tracks control performance, while single-step prediction error shows virtually no correlation.

latent world modelsmodel-based controloff-manifold divergence

Hot Scholars

SZ

Shiduo Zhang

Fudan University
Embodied AIFoundation Models
RQ

Ruihong Qiu

ARC DECRA Fellow, Lecturer (Assistant Professor) @The University of Queensland
GraphLarge Language Models
BG

Bahman Gharesifard

Professor of Mathematics at Queen's University
Control TheoryOptimizationReinforcement LearningNeural Networks
CY

Chengyang Ying

Tsinghua university
Machine LearningReinforcement LearningEmbodied AI
YD

Yi Ding

University of Texas at Dallas
Cyber-Physical SystemsMobile ComputingMachine Learning