action-conditioned modeling

Learning predictive or generative models that produce future states or observations conditional on candidate actions or action sequences, used to imagine outcomes, support planning, and integrate with memory and online decision-making.

action-conditionedmodeling

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

This work addresses the lack of a unified predictive framework for world models in robotic manipulation, which has led to fragmented research and ambiguous design choices. Focusing on three core questions—what to predict, how to relate predictions to actions, and when to use predictions—the paper proposes a functional taxonomy that distinguishes between integrated prediction-action models and explicit predictive planners, positioning world models as foundational predictive infrastructure for robot learning. The study presents a systematic review covering latent dynamics models, action-conditioned video generation, 3D/4D scene prediction, physics simulators, and prediction modules in vision–language–action systems. It also consolidates evaluation protocols across 34 manipulation datasets, highlighting open challenges such as contact modeling and hallucination control through the lenses of prediction fidelity, task performance, and simulation reliability.

action-conditioned predictionbenchmarkingpredictive modeling

Current world models suffer from conceptual ambiguity in embodied intelligence and generative simulation, lacking a unified classification and design framework tailored for robotic control. This work formally defines a world model as one conditioned on actions to predict the future evolution of task-relevant observations or states, and introduces a novel paradigm—world action models—that explicitly links prediction with executable actions. Building on this definition, the study systematically organizes four methodological families: imagine-then-execute, video-feature-conditioned action prediction, joint video-action modeling, and policy learning augmented by auxiliary video prediction. The research clarifies the conceptual boundaries of (action-conditioned) world models and presents the first structured taxonomy specifically designed for embodied prediction and control, thereby advancing standardized understanding and application in the field.

action-conditioned predictionembodied intelligencegenerative simulation

Must-Read Papers

Most classic and influential ideas
View more

Prospective Learning: Learning for a Dynamic Future

Oct 31, 2024
AD
Ashwin De Silva
🏛️ Johns Hopkins University | University of Pennsylvania

Traditional PAC learning theory and static Empirical Risk Minimization (ERM) are fundamentally inadequate for dynamic learning settings where data distributions and task objectives evolve over time. Method: We propose a prospective learning theoretical framework that explicitly incorporates time as an input variable, enabling the construction of sequential predictors. We introduce the first theoretically grounded learning model tailored to time-varying environments—Prospective ERM—and prove that its empirical risk converges to the time-varying Bayesian optimal solution. Contribution/Results: Our analysis formally establishes the inherent failure of static ERM under distributional drift. Leveraging time-varying stochastic process modeling and generalization bounds, experiments on synthetic benchmarks and temporally extended MNIST/CIFAR-10 tasks demonstrate substantial improvements over standard ERM. The framework provides a novel paradigm for robust learning in non-stationary environments, bridging theoretical rigor with empirical efficacy.

ERM MethodPAC Learning TheoryTime-Varying Data

Current agents exhibit underutilization, misinterpretation of predictions, or performance degradation when employing generative world models for prospective reasoning. This work presents the first systematic evaluation of vision-language model–based agents’ ability to strategically invoke and integrate world model predictions across multitask scenarios, combining visual question answering with agent benchmark tasks and introducing attribution analysis to quantitatively measure simulation usage. The study reveals that agents proactively invoke simulations in fewer than 1% of cases, approximately 15% of predictions are misused, and enforced simulation use can degrade performance by up to 5%. These findings expose a cognitive bottleneck in how agents interpret and incorporate predictive information, underscoring the urgent need for calibrated mechanisms to enable reliable prospective reasoning.

agent cognitionanticipatory reasoningforesight

Spatiotemporal Forecasting as Planning: A Model-Based Reinforcement Learning Approach with Generative World Models

Oct 04, 2025
HW
Hao Wu
🏛️ Tsinghua University | OpenAI | Tencent Hunyuan | SLAI | CUHK | Tencent Jarvis Lab | University of Wisconsin | Nanyang Technological University

Addressing the dual challenges of inherent stochasticity and non-differentiable evaluation metrics in physical spatiotemporal forecasting, this paper proposes a novel model-based reinforcement learning paradigm that reformulates prediction as sequential planning. Methodologically, we construct a generative world model to simulate high-fidelity, diverse future states and employ domain-specific non-differentiable metrics—such as extreme-event hit rate—as sparse reward signals. We design a beam-search–guided, reward-driven imagination mechanism and introduce an iterative pseudo-labeling self-training strategy. Crucially, our framework enables end-to-end optimization of non-differentiable objectives without gradient approximation. Experiments demonstrate substantial reductions in overall prediction error alongside marked improvements in long-tail event detection. This work establishes a new pathway toward interpretable and robust forecasting for complex physical systems.

Addresses spatiotemporal forecasting challenges with stochasticity and non-differentiable metricsOptimizes forecasting through planning algorithms using non-differentiable reward signalsProposes model-based reinforcement learning with generative world simulation

This work addresses the absence of a unified closed-loop learning environment that enables agents to continuously learn from real-world events and forecast future outcomes. To bridge this gap, we propose FutureWorld—the first framework that formulates real-time future prediction as a reinforcement learning environment. By integrating a closed-loop mechanism of prediction, outcome realization, and parameter update, FutureWorld effectively prevents answer leakage and supports continual learning. Built upon open-source large language models and grounded in real-world event feedback, the framework establishes a daily-updated benchmark for training and evaluation. Experimental results over consecutive days demonstrate the efficacy of our approach, setting a new state-of-the-art baseline and significantly advancing agents’ predictive capabilities.

agent traininglive future predictionpredictive agents

Generative Models, Humans, Predictive Models: Who Is Worse at High-Stakes Decision Making?

Oct 20, 2024
ST
Sarah Tan
🏛️ Cornell University | University of Washington | Guide Labs | Microsoft Research

This study systematically evaluates the suitability of mainstream large language models (LLMs) for high-stakes judicial recidivism prediction—a domain demanding high accuracy, robustness, and fairness. Method: We conduct a rigorous comparative assessment against human judgments and domain-specific predictive models, employing consistency analysis, adversarial prompt engineering, irrelevant information perturbation (e.g., extraneous photographs), and bias stress testing. Contribution/Results: Our empirical analysis is the first to demonstrate that LLMs underperform domain-specialized models across all dimensions—accuracy, robustness, and fairness—and exhibit significant susceptibility to irrelevant inputs. Critically, several widely adopted bias-mitigation techniques exacerbate decision distortion rather than alleviate it. These findings challenge the viability of deploying LLMs as direct substitutes for human experts or purpose-built models in high-stakes decision-making contexts, providing critical empirical evidence for AI governance and responsible deployment in sensitive domains.

Comparison with human and predictive modelsGenerative models in high-stakes decisionsImpact of information types on decisions

Latest Papers

What's happening recently
View more

This work addresses the computational inefficiency and action latency inherent in traditional embodied control methods that rely on real-time, high-cost world-generation models. The authors propose internalizing the multi-level intermediate states produced by a future-prediction generator into a predictive representation trained via supervised learning, which depends solely on current visual–language inputs. This approach enables efficient control without explicit future generation for the first time. By “folding” the generator’s internal structure into a present-moment representation, the method supports rapid adaptation to dynamic environmental perturbations and human interventions. Evaluated on LIBERO, RoboTwin2.0, and real-robot tasks, the framework reduces inference latency by 3.7–10.1×, effectively suppresses irrelevant scene variations, captures long-horizon dynamics, and demonstrates robust flexibility under manual intervention.

action latencyefficient inferenceembodied control

This work proposes the Imagine-then-Plan (ITP) framework to address the limitations of existing world models, which typically rely on single-step or fixed-horizon rollouts and struggle to support effective planning in complex tasks. ITP enables the policy to interact with the world model to generate multi-step imagined trajectories and introduces a novel adaptive lookahead mechanism that dynamically adjusts the imagination horizon to integrate current observations with future predictions, thereby guiding policy learning. By unifying multi-step imagination with partially observable Markov decision processes, the approach supports both zero-shot and reinforcement learning paradigms. Extensive experiments across multiple agent benchmarks demonstrate that ITP significantly outperforms current methods, validating the effectiveness of the adaptive lookahead mechanism in enhancing task completion performance and reasoning depth.

adaptive lookaheadagent planningcomplex task planning

This work addresses the limitation of existing end-to-end autonomous driving approaches, which are predominantly reactive and lack proactive anticipation of future states. To overcome this, the authors propose ForeSight, a framework that reframes motion planning as a forward-looking decision-making process. ForeSight leverages a pretrained world model to generate multimodal future visual scenes, which in turn conditionally guide action planning—elevating the world model from an auxiliary component to the core planning backbone. Notably, this is the first method to explicitly incorporate imagined future scenarios as the central mechanism for action prediction. Evaluated on the NAVSIM and nuScenes datasets, ForeSight significantly outperforms current state-of-the-art methods, demonstrating the efficacy and superiority of prospective modeling in complex, dynamic traffic environments.

anticipatory decision-makingautonomous drivingforesight

This work addresses the limitations of current world models, which often prioritize visual fidelity at the expense of physical plausibility and causal structure, thereby hindering their capacity for intervention, long-horizon prediction, and safety-critical decision-making. To overcome these shortcomings, the paper proposes a novel paradigm grounded in physical realism and explicit causal modeling, reframing world models as actionable simulators. The approach integrates a structured 4D interface, constraint-aware dynamics, and counterfactual reasoning mechanisms to enable precise intervention planning. Furthermore, it introduces a closed-loop evaluation framework to rigorously assess model performance. Evaluated in high-stakes domains such as medical decision-making, the method demonstrates substantial improvements in long-term robustness, intervention efficacy, and causal consistency.

actionable simulatorscausal dynamicsphysical grounding

Existing vision-language-action (VLA) models struggle to effectively internalize physical world knowledge, as their video prediction often degenerates into simplistic extrapolation and fails to jointly capture instantaneous dynamics and long-term causal dependencies. This work proposes a VLA architecture embedded with a predictive world model that leverages a chunked autoregressive mechanism for efficient long-horizon causal forecasting. It introduces a temporal importance sampling strategy grounded in egomotion and behavioral signals, combined with a curriculum-based progressive training paradigm. Furthermore, a diffusion-based multi-view renderer is integrated to enhance visual fidelity. The resulting approach substantially improves long-term planning capabilities while preserving high-fidelity visual generation, establishing a new paradigm for knowledge-driven autonomous agents.

long-horizon causalityphysical world knowledgepredictive world modeling

Hot Scholars

SL

Sergey Levine

UC Berkeley, Physical Intelligence
Machine LearningRoboticsReinforcement Learning
YD

Yilun Du

Harvard University
Artificial IntelligenceMachine LearningRoboticsComputer Vision
PA

Pieter Abbeel

UC Berkeley | Covariant
RoboticsMachine LearningAI
YZ

Yuke Zhu

The University of Texas at Austin, NVIDIA Research
Robot LearningComputer VisionMachine LearningRobotics