Score
Design, build, and evaluate models and inference procedures that predict or generate future visual observations (video frames or multi-view imagery) conditioned on action sequences and past observations, producing temporally consistent, high‑fidelity rollouts together with corresponding action or proprioceptive trajectories and control signals. This work includes constructing dynamic world and world‑action models and executors, implementing mixed autoregressive–bidirectional decoding and hybrid bidirectional–autoregressive inference (including converting bidirectional backbones to autoregressive use), and producing self‑correcting long‑horizon simulated futures usable for planning, policy training, or action‑conditioned control.
This study addresses the conceptual ambiguity surrounding World Action Models (WAMs) by clarifying their distinctions from related paradigms such as world models, video generation, and vision-language action policies. Through two complementary lenses—generated content type and methodological composition—the work proposes the first unified taxonomy to systematically deconstruct WAM design paradigms. It reveals a fundamental trade-off between representational richness and computational, memory, latency, and action annotation costs, highlighting a growing trend toward generating minimal yet control-critical future information. The paper characterizes WAMs as inherently predictive-action synergistic mechanisms, distills common design patterns, and provides a systematic overview of current advances and open challenges across key dimensions including interactivity, causality, persistence, physical plausibility, and generalization capability.
This work addresses the rapid degradation of visual fidelity in autoregressive multi-step rollouts of action-conditioned robotic world models, a problem primarily caused by error accumulation. To mitigate this, the authors propose a post-training framework grounded in contrastive reinforcement learning, which optimizes the model using its own generated rollout sequences. The approach introduces a multi-candidate, variable-length future comparison mechanism and a visual fidelity reward that integrates multi-view perceptual metrics. This design significantly enhances both consistency and realism in long-horizon predictions. Evaluated on the DROID dataset, the method achieves state-of-the-art performance in rollout fidelity, reducing LPIPS by 14% and improving SSIM by 9.1%, with human evaluators expressing an 80% preference for its outputs in blind tests.
Current long-video generation models suffer from compounding errors and weak memory mechanisms, hindering the construction of foundational world models that simultaneously ensure interactivity and spatiotemporal consistency. To address this, we propose Video Retrieval-Augmented Generation (VRAG), the first framework to introduce explicit global state conditioning—overcoming the inherent, irreducible error accumulation bottleneck in autoregressive modeling. VRAG integrates action-conditioned generation with efficient video retrieval to substantially suppress long-horizon generation errors. Furthermore, we establish the first comprehensive benchmark specifically designed to evaluate world-modeling capabilities in video generation. Extensive experiments demonstrate significant improvements in spatiotemporal consistency, interactive controllability, and long-sequence fidelity. Our approach establishes a novel paradigm for interactive, video-based world models.
Existing pre-trained video generation models rely on static prompts (e.g., text or images), limiting their ability to model interactive, dynamic scenes. To address this, we propose the Dynamic World Simulation (DWS) framework, which transforms video generation models into interactive world simulators. DWS introduces a lightweight, universal action-conditioning module that drives scene evolution according to given action trajectories; a motion-augmented loss that explicitly optimizes dynamic consistency—rather than pixel-level fidelity; and a priority imagination sampling strategy to enhance long-horizon temporal controllability. The framework is architecture-agnostic, supporting both diffusion models and autoregressive Transformers. Experiments demonstrate that DWS significantly improves action controllability and dynamic coherence in game and robotics simulation scenarios. Moreover, when applied to downstream model-predictive control tasks, DWS achieves state-of-the-art sample efficiency.
Existing autoregressive world models suffer from spatial structural distortion, low decoding efficiency, and weak motion modeling in video prediction. To address these issues, we propose a generative world model that establishes a hybrid spatiotemporal modeling paradigm: it integrates intra-frame bidirectional spatial attention with causal temporal decoding, introduces a trajectory-aware motion prompting module, and employs an asymmetric multi-scale tokenizer—while enabling parallel autoregressive decoding. Our framework significantly improves spatiotemporal consistency and physical plausibility. It achieves state-of-the-art performance on action-conditioned video prediction and model-based control tasks. Moreover, inference speed is accelerated by 4.4× compared to baseline methods. The model demonstrates zero-shot transfer capability across domains and exhibits strong scalability to varying input resolutions and sequence lengths.
This work proposes LingBot-VA, an autoregressive diffusion-based control framework that integrates video world modeling with causal reasoning to enhance long-horizon robotic control and generalization in complex environments. By leveraging a Mixture-of-Transformers architecture, the method constructs a shared latent space for vision and action, enabling joint learning of video frame prediction and policy execution. It further incorporates closed-loop rolling inference and asynchronous parallel control mechanisms to improve temporal coherence and responsiveness. Experimental results demonstrate that LingBot-VA significantly outperforms baseline approaches in both simulation and real-world settings, achieving higher success rates on long-horizon tasks, improved data efficiency, and stronger generalization to novel environmental configurations.
This work addresses the limited action sensitivity and poor controllability of action-conditioned world models, which often rely on statistical shortcuts such as visual inertia. To enhance action controllability, we propose the CoCo framework, which enforces multi-step and action–spatial counterfactual consistency constraints to ensure prediction invariance under perturbations to actions or observations. We introduce evaluation protocols—including the ARC, DE, and Mini-SSMB benchmarks—as well as mirror-world environments, action transformations, and quantitative controllability metrics. Experiments demonstrate that CoCo achieves an ARC_inv of 0.412 and ARC_ref of 0.483 on Mini-SSMB, reduces drift energy by 17.07%, and attains a 73.1% success rate on the VP2 visual planning task, outperforming current state-of-the-art methods.
Current world models suffer from conceptual ambiguity in embodied intelligence and generative simulation, lacking a unified classification and design framework tailored for robotic control. This work formally defines a world model as one conditioned on actions to predict the future evolution of task-relevant observations or states, and introduces a novel paradigm—world action models—that explicitly links prediction with executable actions. Building on this definition, the study systematically organizes four methodological families: imagine-then-execute, video-feature-conditioned action prediction, joint video-action modeling, and policy learning augmented by auxiliary video prediction. The research clarifies the conceptual boundaries of (action-conditioned) world models and presents the first structured taxonomy specifically designed for embodied prediction and control, thereby advancing standardized understanding and application in the field.
This work addresses the challenge of effectively modeling counterfactual outcomes following action interventions in interactive video world models. The authors propose a noise-coupled dual-branch unrolling mechanism that, building upon a shared state prefix and exogenous noise, bifurcates only the action stream after the intervention point to enable precise counterfactual generation. By explicitly recovering exogenous noise from self-generated trajectories, the method circumvents the traditional difficulties of approximate inversion and reformulates the principle of minimal change into a verifiable spatiotemporal locality metric. Grounded in Pearl’s causal framework, the approach integrates state branching, noise coupling, and computable causal descendant regions to construct a discriminator-free counterfactual evaluation system, thereby providing reinforcement learning with reliable reward signals.
This work addresses the limitation of existing end-to-end vision-and-language navigation (VLN) methods, which supervise only the current action and lack explicit modeling of future states. To overcome this, the authors propose the FSC-VLN framework, which introduces a future state conditioning mechanism: during training, future visual embeddings serve as supervisory signals to guide the policy toward learning forward-looking state representations, while requiring no future images at inference time. Built upon a causal vision-language model, FSC-VLN features a dual-branch architecture with separate future-query and action-query streams, employs a frozen visual encoder, and aligns future latent states through a dedicated target branch. The approach achieves significant improvements on the R2R val-unseen split across success rate (SR), oracle success rate (OSR), and success weighted by path length (SPL), with particularly strong performance on long-horizon trajectories.