predictive world modeling

Designs, builds, or analyzes internal world models that represent entities, their attributes, and their spatiotemporal dynamics using structured, schema- or symbol-guided, probabilistic, or learned transition representations; these artifacts include belief-based memories, parameterized/DSGE-style structured models, physics-grounded symbolic architectures, and probabilistic scene models. Such work implements and evaluates multi-step mental rollouts and future-state prediction, uncertainty-aware belief updating (e.g., closed-form Bayesian posteriors, entropy-driven decay), information-theoretic scoring of observations, and learned transition functions that support decision-making and planning.

predictiveworldmodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.08
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$235K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

This work proposes a unified framework to address the integrated demands of prediction, reasoning, and decision-making in real-world applications such as robotics and autonomous driving. By systematically combining explicit and implicit world modeling paradigms—categorized and fused according to their representational structures and utilization in prediction—the framework clarifies the theoretical foundations of physical AI within the perception–prediction–action loop. It further delineates key challenges including hierarchical reasoning, long-horizon planning, and autonomous goal generation, thereby offering a coherent pathway toward artificial general intelligence and enabling intelligent systems to evolve from reactive control to anticipatory decision-making.

autonomous goal formationhierarchical reasoninglong-horizon planning

From 2D to 3D Cognition: A Brief Survey of General World Models

Jun 25, 2025
NX
Ningwei Xie
🏛️ China Mobile Research Institute | Shenzhen Ubiquitous Data Enabling Key Lab | Shenzhen International Graduate School | Tsinghua University

A systematic taxonomy and survey of world models bridging 2D perception to 3D cognition remains absent. Method: We propose a dual-axis classification framework—“3D representation advancement” and “world knowledge integration”—to systematically characterize the technical evolution of 3D-cognitive world models. Our approach unifies neural radiance fields (NeRF), 3D generative modeling, spatial reasoning networks, physics simulation, and multimodal knowledge integration into a coherent conceptual framework for 3D spatial cognition. Contribution/Results: We formally identify three core capabilities—3D scene generation, spatial reasoning, and embodied interaction—and clarify their interdependencies. Furthermore, we pinpoint critical challenges including data scarcity, limited modeling generalizability, and real-time deployment constraints. This work establishes a foundational theoretical basis and provides a principled technical roadmap toward developing generalizable, robust 3D-cognitive systems.

Challenges in data, modeling, and deployment of 3D modelsLack of systematic analysis for 3D cognitive world modelsTransition from 2D perception to 3D cognition in AI

Must-Read Papers

Most classic and influential ideas
View more

Learning World Models With Hierarchical Temporal Abstractions: A Probabilistic Perspective

Apr 24, 2024
VS
Vaisakh Shaj Kumar
🏛️ Karlsruher Institut für Technologie (KIT)

Existing state space models (SSMs) struggle to capture non-stationary, multi-scale causal dynamics prevalent in real-world systems. Method: We propose a scalable hierarchical world model that jointly integrates latent-parameter SSMs with multi-timescale SSMs, enabling— for the first time—graph-structured, exact probabilistic inference with end-to-end temporal learning. Leveraging probabilistic graphical models, belief propagation, and Bayesian inference, our approach explicitly represents uncertainty to better approximate inherent stochasticity. Contribution/Results: Evaluated on diverse real and simulated robotic tasks, the model matches or surpasses leading Transformer variants in long-horizon future prediction while substantially enhancing cross-temporal and cross-spatial causal reasoning capabilities.

Addressing limitations of state space models with new formalismsDeveloping hierarchical world models for multi-level reasoningIntegrating uncertainty to improve real-world dynamics representation

World Modeling with Probabilistic Structure Integration

Sep 10, 2025
KK
Klemen Kotar
🏛️ Stanford NeuroAI Lab

Existing world models lack both strong controllability and flexible prompting capabilities for structured scene understanding. Method: We propose a “probabilistic prediction–structural extraction–integrated optimization” three-stage iterative learning framework. First, zero-shot causal inference disentangles implicit intermediate representations (e.g., optical flow, depth, semantic segmentation) from raw video data; these are then encoded as novel learnable tokens integrated into a unified, LLM-inspired prompting architecture. Technically, the framework synergistically combines probabilistic graphical models, stochastic autoregressive modeling with random access, causal inference, and self-supervised learning. Contribution/Results: Evaluated on trillion-frame video datasets, our model achieves state-of-the-art performance across multiple vision tasks—including optical flow estimation, monocular depth prediction, and object segmentation—while enabling cross-task prompt-based control and continual performance improvement. To our knowledge, this is the first work to unify structured world modeling with general-purpose, instruction-tunable prompting mechanisms.

Extracting low-dimensional structures via causal inferenceImproving video prediction and understanding capabilitiesLearning controllable world models from video data

This work addresses the challenge of modeling complex action effects and causal relationships in data-scarce yet semantically rich real-world domains, such as commercial environments, where effective planning is critical. The authors propose CASSANDRA, a neuro-symbolic world modeling approach that leverages large language models to provide knowledge priors, guiding the generation of procedural transition rules. These rules are integrated with probabilistic graphical models for structural learning, enabling joint modeling of deterministic action effects and stochastic variable dependencies. Evaluated in simulated coffee shop and theme park environments, CASSANDRA significantly outperforms existing baselines, achieving marked improvements in both state transition prediction accuracy and success rates on downstream planning tasks.

causal relationshipslimited dataplanning

Natural Building Blocks for Structured World Models: Theory, Evidence, and Scaling

Nov 03, 2025
LD
Lancelot Da Costa
🏛️ VERSES AI Research Lab | University of Tübingen | ELLIS Institute

The field of world modeling has long suffered from a lack of unified theoretical foundations, resulting in fragmented architectures and poor interpretability. Method: This paper proposes a principled framework for constructing structured world models, unifying discrete (logical/symbolic) and continuous (physical/dynamical) stochastic processes as core modeling paradigms. It hierarchically composes Hidden Markov Models (HMMs) with switching Linear Dynamical Systems (sLDS), enforcing a fixed causal structure while optimizing only depth parameters—thereby avoiding combinatorial explosion. The framework integrates Partially Observable Markov Decision Processes (POMDPs) with controllable sLDS and employs an incremental joint learning strategy for structure and parameters. Contribution/Results: Evaluated on multimodal generation and pixel-level planning tasks, the model matches deep neural networks in performance while retaining explicit semantic grounding and full traceability—demonstrating the effectiveness, modularity, and interpretability of the proposed paradigm.

Addresses combinatorial explosion in structure learning by fixing causal architecture with minimal parametersProposes fundamental building blocks for structured world models combining discrete and continuous processesSolves scalable joint structure-parameter learning challenge while maintaining model interpretability

Existing world models struggle to infer the complete physical structure of scenes and interactions among objects from partially observed videos. This work proposes a novel probabilistic world model based on autoregressive sequence modeling, which enables efficient training and supports conditional estimation over arbitrary visual variables—such as appearance and dynamics. The model generates multimodal future states over multiple steps, automatically discovers objects and their subparts, and facilitates 3D manipulation and physical reasoning. Experiments demonstrate that the model successfully extracts hierarchical object structures in tasks such as Visual Jenga, significantly enhancing the understanding of complex physical interactions.

object interactionphysical object understandingscene structure

Latest Papers

What's happening recently
View more

This work addresses the challenge of modeling structured uncertainty for embodied agents in partially observable environments—a limitation of current vision-based generative models that prioritize photorealism. The paper formulates world modeling as embodied belief inference in 3D space and introduces the first generative world model that explicitly represents uncertainty directly in 3D. This approach enables spatially consistent scene memory, multi-hypothesis belief sampling, temporal belief updates, and semantics-guided prediction of unobserved regions. By integrating multi-view geometry, probabilistic reasoning, and semantic priors, the method supports online inference and updating of 3D beliefs. It outperforms existing approaches in both 2D/3D scene reconstruction quality and downstream embodied tasks such as object navigation, with validation in both simulated and real-world environments.

3D uncertaintyembodied belief inferencepartial observability

Existing world models struggle to explicitly represent entity attributes, interaction relations, and causal structures in the environment, limiting the reliability of prediction and decision-making. This work proposes a Causal World Model (CWM) that integrates causal representation learning, object-centric modeling, structural causal models, and causal discovery into a unified theoretical framework endowed with both generative capacity and interpretable causal mechanisms. The framework bridges perception, conceptual representation, and dynamics modeling, clarifies the model’s role in decision tasks, and characterizes the identifiability conditions and equivalence classes recoverable from observational data, thereby establishing a theoretical foundation for robust causal reasoning and decision-making.

causal reasoningcausal representation learningCausal World Models

This work addresses the limitation of existing world models, which focus solely on physical states and thus struggle to accurately predict human-driven behaviors. To overcome this, the paper introduces the first world modeling framework that explicitly incorporates mental variables—such as beliefs and intentions—into its core architecture. It proposes a physics–mind coupled dynamic mechanism to jointly model their interactions and integrates a modular pipeline comprising state parsing, goal observation generation, action decomposition, joint transition modeling, and branching value evaluation, leveraging large language models for interpretable mental-state reasoning. The authors also provide MENTIS, a trainable-free, testable baseline, and demonstrate through experiments on multimodal situated decision-making datasets that explicit modeling of mental states is crucial for behavior prediction, thereby revealing key bottlenecks in current approaches to mental modeling.

belief modelinghuman decision predictionmental state

This work proposes a physics-informed world model that addresses the limitations of conventional video prediction approaches, which typically model dynamics directly in pixel space and struggle to capture underlying physical principles explicitly. The proposed method learns a compact, discrete "physics language" from unlabeled real-world videos through self-supervision, enabling explicit representation of world states. It adopts a "reason-then-render" paradigm: future states are predicted by performing action-conditioned sequential reasoning in this discrete latent space, followed by rendering to generate future video frames. This approach yields interpretable, physically consistent dynamics, demonstrating strong performance in both generative and perceptual tasks. Moreover, it supports interactive simulation, fine-grained action control, and zero-shot motion transfer, highlighting its capacity for structured and generalizable physical reasoning.

discrete representationexplicit reasoningphysical world modeling

This work addresses the challenge of unifying prediction, planning, and irreversibility within world models. It formulates prediction as a probability measure over future trajectories and, under a local Markov assumption, employs the Onsager–Machlup action functional to decompose latent dynamics into reversible and irreversible components in path space. The authors introduce rollout-based entropy production as an operational measure of irreversibility. Through path-integral analysis, attention mechanism inspection, and small-scale model experiments, they find that attention asymmetry emerges in response to increasing data irreversibility. While symmetrization interventions suppress entropy production, they selectively impair long-horizon prediction of irreversible processes yet preserve the model’s capacity to capture relaxation dynamics.

entropy productionirreversibilitypath-space

Hot Scholars

XL

Xiu Li

Bytedance Seed
Computer VisionComputer Graphics3D Vision
XC

Xiaowei Chi

The Hong Kong University of Science and Technology
Multimodal GenerationRoboticsComputer Vision
TA

Tim Althoff

Associate Professor of Computer Science, University of Washington
Human AI InteractionNatural Language ProcessingBehavioral Data ScienceAI for Mental Health
MY

Mingqi Yuan

PhD candidate at HKPU
Machine Learning