action-conditional video prediction

Design, build, and evaluate models and inference procedures that predict or generate future visual observations (video frames or multi-view imagery) conditioned on action sequences and past observations, producing temporally consistent, high‑fidelity rollouts together with corresponding action or proprioceptive trajectories and control signals. This work includes constructing dynamic world and world‑action models and executors, implementing mixed autoregressive–bidirectional decoding and hybrid bidirectional–autoregressive inference (including converting bidirectional backbones to autoregressive use), and producing self‑correcting long‑horizon simulated futures usable for planning, policy training, or action‑conditioned control.

action-conditionalvideoprediction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.32
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the rapid degradation of visual fidelity in autoregressive multi-step rollouts of action-conditioned robotic world models, a problem primarily caused by error accumulation. To mitigate this, the authors propose a post-training framework grounded in contrastive reinforcement learning, which optimizes the model using its own generated rollout sequences. The approach introduces a multi-candidate, variable-length future comparison mechanism and a visual fidelity reward that integrates multi-view perceptual metrics. This design significantly enhances both consistency and realism in long-horizon predictions. Evaluated on the DROID dataset, the method achieves state-of-the-art performance in rollout fidelity, reducing LPIPS by 14% and improving SSIM by 9.1%, with human evaluators expressing an 80% preference for its outputs in blind tests.

autoregressive predictionerror compoundingmulti-step rollouts

Learning World Models for Interactive Video Generation

May 28, 2025
TC
Taiye Chen
🏛️ Peking University | University of Oxford | Princeton University

Current long-video generation models suffer from compounding errors and weak memory mechanisms, hindering the construction of foundational world models that simultaneously ensure interactivity and spatiotemporal consistency. To address this, we propose Video Retrieval-Augmented Generation (VRAG), the first framework to introduce explicit global state conditioning—overcoming the inherent, irreducible error accumulation bottleneck in autoregressive modeling. VRAG integrates action-conditioned generation with efficient video retrieval to substantially suppress long-horizon generation errors. Furthermore, we establish the first comprehensive benchmark specifically designed to evaluate world-modeling capabilities in video generation. Extensive experiments demonstrate significant improvements in spatiotemporal consistency, interactive controllability, and long-sequence fidelity. Our approach establishes a novel paradigm for interactive, video-based world models.

Addressing compounding errors in autoregressive video generationEnhancing video models with interactive action conditioningImproving spatiotemporal coherence with retrieval-augmented generation

Pre-Trained Video Generative Models as World Simulators

Feb 10, 2025
HH
Haoran He
🏛️ Hong Kong University of Science and Technology | Tsinghua University | Sun Yat-sen University | Tencent AI Lab

Existing pre-trained video generation models rely on static prompts (e.g., text or images), limiting their ability to model interactive, dynamic scenes. To address this, we propose the Dynamic World Simulation (DWS) framework, which transforms video generation models into interactive world simulators. DWS introduces a lightweight, universal action-conditioning module that drives scene evolution according to given action trajectories; a motion-augmented loss that explicitly optimizes dynamic consistency—rather than pixel-level fidelity; and a priority imagination sampling strategy to enhance long-horizon temporal controllability. The framework is architecture-agnostic, supporting both diffusion models and autoregressive Transformers. Experiments demonstrate that DWS significantly improves action controllability and dynamic coherence in game and robotics simulation scenarios. Moreover, when applied to downstream model-predictive control tasks, DWS achieves state-of-the-art sample efficiency.

Action-controllable video generationDynamic world simulationEnhancing video generative models

SAMPO:Scale-wise Autoregression with Motion PrOmpt for generative world models

Sep 18, 2025
SW
Sen Wang
🏛️ Xi’an Jiaotong University | University of Illinois at Chicago | Amazon.com, Inc.

Existing autoregressive world models suffer from spatial structural distortion, low decoding efficiency, and weak motion modeling in video prediction. To address these issues, we propose a generative world model that establishes a hybrid spatiotemporal modeling paradigm: it integrates intra-frame bidirectional spatial attention with causal temporal decoding, introduces a trajectory-aware motion prompting module, and employs an asymmetric multi-scale tokenizer—while enabling parallel autoregressive decoding. Our framework significantly improves spatiotemporal consistency and physical plausibility. It achieves state-of-the-art performance on action-conditioned video prediction and model-based control tasks. Moreover, inference speed is accelerated by 4.4× compared to baseline methods. The model demonstrates zero-shot transfer capability across domains and exhibits strong scalability to varying input resolutions and sequence lengths.

Addressing spatial structure disruption and inefficient decoding issuesEnhancing motion modeling and dynamic scene understanding efficiencyImproving visual coherence in autoregressive world model predictions

This work proposes LingBot-VA, an autoregressive diffusion-based control framework that integrates video world modeling with causal reasoning to enhance long-horizon robotic control and generalization in complex environments. By leveraging a Mixture-of-Transformers architecture, the method constructs a shared latent space for vision and action, enabling joint learning of video frame prediction and policy execution. It further incorporates closed-loop rolling inference and asynchronous parallel control mechanisms to improve temporal coherence and responsiveness. Experimental results demonstrate that LingBot-VA significantly outperforms baseline approaches in both simulation and real-world settings, achieving higher success rates on long-horizon tasks, improved data efficiency, and stronger generalization to novel environmental configurations.

action-visual dynamicscausal world modelinglong-horizon manipulation

Latest Papers

What's happening recently
View more

This work addresses the limited action sensitivity and poor controllability of action-conditioned world models, which often rely on statistical shortcuts such as visual inertia. To enhance action controllability, we propose the CoCo framework, which enforces multi-step and action–spatial counterfactual consistency constraints to ensure prediction invariance under perturbations to actions or observations. We introduce evaluation protocols—including the ARC, DE, and Mini-SSMB benchmarks—as well as mirror-world environments, action transformations, and quantitative controllability metrics. Experiments demonstrate that CoCo achieves an ARC_inv of 0.412 and ARC_ref of 0.483 on Mini-SSMB, reduces drift energy by 17.07%, and attains a 73.1% success rate on the VP2 visual planning task, outperforming current state-of-the-art methods.

action controllabilityaction-conditioned world modelscounterfactual consistency

Current world models suffer from conceptual ambiguity in embodied intelligence and generative simulation, lacking a unified classification and design framework tailored for robotic control. This work formally defines a world model as one conditioned on actions to predict the future evolution of task-relevant observations or states, and introduces a novel paradigm—world action models—that explicitly links prediction with executable actions. Building on this definition, the study systematically organizes four methodological families: imagine-then-execute, video-feature-conditioned action prediction, joint video-action modeling, and policy learning augmented by auxiliary video prediction. The research clarifies the conceptual boundaries of (action-conditioned) world models and presents the first structured taxonomy specifically designed for embodied prediction and control, thereby advancing standardized understanding and application in the field.

action-conditioned predictionembodied intelligencegenerative simulation

This work addresses the challenge of effectively modeling counterfactual outcomes following action interventions in interactive video world models. The authors propose a noise-coupled dual-branch unrolling mechanism that, building upon a shared state prefix and exogenous noise, bifurcates only the action stream after the intervention point to enable precise counterfactual generation. By explicitly recovering exogenous noise from self-generated trajectories, the method circumvents the traditional difficulties of approximate inversion and reformulates the principle of minimal change into a verifiable spatiotemporal locality metric. Grounded in Pearl’s causal framework, the approach integrates state branching, noise coupling, and computable causal descendant regions to construct a discriminator-free counterfactual evaluation system, thereby providing reinforcement learning with reliable reward signals.

causal interventioncounterfactual generationinteractive video world models

This work addresses the limitation of existing end-to-end vision-and-language navigation (VLN) methods, which supervise only the current action and lack explicit modeling of future states. To overcome this, the authors propose the FSC-VLN framework, which introduces a future state conditioning mechanism: during training, future visual embeddings serve as supervisory signals to guide the policy toward learning forward-looking state representations, while requiring no future images at inference time. Built upon a causal vision-language model, FSC-VLN features a dual-branch architecture with separate future-query and action-query streams, employs a frozen visual encoder, and aligns future latent states through a dedicated target branch. The approach achieves significant improvements on the R2R val-unseen split across success rate (SR), oracle success rate (OSR), and success weighted by path length (SPL), with particularly strong performance on long-horizon trajectories.

actionable cuesbehavior cloningfuture-state prediction

Hot Scholars

DF

Dieter Fox

University of Washington and AI2
RoboticsArtificial IntelligenceComputer Vision
PL

Percy Liang

Associate Professor of Computer Science, Stanford University
machine learningnatural language processing
YF

Yanwei Fu

Fudan University
Computer visionmachine learningMultimedia
MT

Masayoshi Tomizuka

Mechaniccal Engineering, University of California
mechanical engineeringdynamic systemscontrolmechatronics
JZ

Jiazhao Zhang

Peking University
Embodied AINavigation3D Vision