Score
Designs, builds, or analyzes internal world models that represent entities, their attributes, and their spatiotemporal dynamics using structured, schema- or symbol-guided, probabilistic, or learned transition representations; these artifacts include belief-based memories, parameterized/DSGE-style structured models, physics-grounded symbolic architectures, and probabilistic scene models. Such work implements and evaluates multi-step mental rollouts and future-state prediction, uncertainty-aware belief updating (e.g., closed-form Bayesian posteriors, entropy-driven decay), information-theoretic scoring of observations, and learned transition functions that support decision-making and planning.
This work addresses the lack of a unified framework in world model research, which has hindered systematic integration across diverse architectures, methodologies, reasoning paradigms, and applications. To bridge this gap, the paper proposes a four-dimensional taxonomy encompassing architectural designs (e.g., state-space models, Transformers, diffusion models, physics-informed networks), methodological families (e.g., language-augmented multimodal systems), reasoning paradigms (e.g., imagination-based planning and counterfactual reasoning), and application domains. For the first time, this framework connects foundational insights from cognitive science to the latest advances in large-scale models, revealing an emerging trend toward integrating chain-of-thought reasoning with world model imagination. It clarifies the overall developmental landscape, identifies core challenges such as error accumulation in prediction and simulation-to-reality transfer, and outlines a pathway toward unified multimodal world models.
This work proposes a unified framework to address the integrated demands of prediction, reasoning, and decision-making in real-world applications such as robotics and autonomous driving. By systematically combining explicit and implicit world modeling paradigms—categorized and fused according to their representational structures and utilization in prediction—the framework clarifies the theoretical foundations of physical AI within the perception–prediction–action loop. It further delineates key challenges including hierarchical reasoning, long-horizon planning, and autonomous goal generation, thereby offering a coherent pathway toward artificial general intelligence and enabling intelligent systems to evolve from reactive control to anticipatory decision-making.
A systematic taxonomy and survey of world models bridging 2D perception to 3D cognition remains absent. Method: We propose a dual-axis classification framework—“3D representation advancement” and “world knowledge integration”—to systematically characterize the technical evolution of 3D-cognitive world models. Our approach unifies neural radiance fields (NeRF), 3D generative modeling, spatial reasoning networks, physics simulation, and multimodal knowledge integration into a coherent conceptual framework for 3D spatial cognition. Contribution/Results: We formally identify three core capabilities—3D scene generation, spatial reasoning, and embodied interaction—and clarify their interdependencies. Furthermore, we pinpoint critical challenges including data scarcity, limited modeling generalizability, and real-time deployment constraints. This work establishes a foundational theoretical basis and provides a principled technical roadmap toward developing generalizable, robust 3D-cognitive systems.
Existing state space models (SSMs) struggle to capture non-stationary, multi-scale causal dynamics prevalent in real-world systems. Method: We propose a scalable hierarchical world model that jointly integrates latent-parameter SSMs with multi-timescale SSMs, enabling— for the first time—graph-structured, exact probabilistic inference with end-to-end temporal learning. Leveraging probabilistic graphical models, belief propagation, and Bayesian inference, our approach explicitly represents uncertainty to better approximate inherent stochasticity. Contribution/Results: Evaluated on diverse real and simulated robotic tasks, the model matches or surpasses leading Transformer variants in long-horizon future prediction while substantially enhancing cross-temporal and cross-spatial causal reasoning capabilities.
Existing world models lack both strong controllability and flexible prompting capabilities for structured scene understanding. Method: We propose a “probabilistic prediction–structural extraction–integrated optimization” three-stage iterative learning framework. First, zero-shot causal inference disentangles implicit intermediate representations (e.g., optical flow, depth, semantic segmentation) from raw video data; these are then encoded as novel learnable tokens integrated into a unified, LLM-inspired prompting architecture. Technically, the framework synergistically combines probabilistic graphical models, stochastic autoregressive modeling with random access, causal inference, and self-supervised learning. Contribution/Results: Evaluated on trillion-frame video datasets, our model achieves state-of-the-art performance across multiple vision tasks—including optical flow estimation, monocular depth prediction, and object segmentation—while enabling cross-task prompt-based control and continual performance improvement. To our knowledge, this is the first work to unify structured world modeling with general-purpose, instruction-tunable prompting mechanisms.
This work addresses the challenge of modeling complex action effects and causal relationships in data-scarce yet semantically rich real-world domains, such as commercial environments, where effective planning is critical. The authors propose CASSANDRA, a neuro-symbolic world modeling approach that leverages large language models to provide knowledge priors, guiding the generation of procedural transition rules. These rules are integrated with probabilistic graphical models for structural learning, enabling joint modeling of deterministic action effects and stochastic variable dependencies. Evaluated in simulated coffee shop and theme park environments, CASSANDRA significantly outperforms existing baselines, achieving marked improvements in both state transition prediction accuracy and success rates on downstream planning tasks.
The field of world modeling has long suffered from a lack of unified theoretical foundations, resulting in fragmented architectures and poor interpretability. Method: This paper proposes a principled framework for constructing structured world models, unifying discrete (logical/symbolic) and continuous (physical/dynamical) stochastic processes as core modeling paradigms. It hierarchically composes Hidden Markov Models (HMMs) with switching Linear Dynamical Systems (sLDS), enforcing a fixed causal structure while optimizing only depth parameters—thereby avoiding combinatorial explosion. The framework integrates Partially Observable Markov Decision Processes (POMDPs) with controllable sLDS and employs an incremental joint learning strategy for structure and parameters. Contribution/Results: Evaluated on multimodal generation and pixel-level planning tasks, the model matches deep neural networks in performance while retaining explicit semantic grounding and full traceability—demonstrating the effectiveness, modularity, and interpretability of the proposed paradigm.
Existing world models struggle to infer the complete physical structure of scenes and interactions among objects from partially observed videos. This work proposes a novel probabilistic world model based on autoregressive sequence modeling, which enables efficient training and supports conditional estimation over arbitrary visual variables—such as appearance and dynamics. The model generates multimodal future states over multiple steps, automatically discovers objects and their subparts, and facilitates 3D manipulation and physical reasoning. Experiments demonstrate that the model successfully extracts hierarchical object structures in tasks such as Visual Jenga, significantly enhancing the understanding of complex physical interactions.
This work addresses the challenge of modeling structured uncertainty for embodied agents in partially observable environments—a limitation of current vision-based generative models that prioritize photorealism. The paper formulates world modeling as embodied belief inference in 3D space and introduces the first generative world model that explicitly represents uncertainty directly in 3D. This approach enables spatially consistent scene memory, multi-hypothesis belief sampling, temporal belief updates, and semantics-guided prediction of unobserved regions. By integrating multi-view geometry, probabilistic reasoning, and semantic priors, the method supports online inference and updating of 3D beliefs. It outperforms existing approaches in both 2D/3D scene reconstruction quality and downstream embodied tasks such as object navigation, with validation in both simulated and real-world environments.
Existing world models struggle to explicitly represent entity attributes, interaction relations, and causal structures in the environment, limiting the reliability of prediction and decision-making. This work proposes a Causal World Model (CWM) that integrates causal representation learning, object-centric modeling, structural causal models, and causal discovery into a unified theoretical framework endowed with both generative capacity and interpretable causal mechanisms. The framework bridges perception, conceptual representation, and dynamics modeling, clarifies the model’s role in decision tasks, and characterizes the identifiability conditions and equivalence classes recoverable from observational data, thereby establishing a theoretical foundation for robust causal reasoning and decision-making.
This work addresses the limitation of existing world models, which focus solely on physical states and thus struggle to accurately predict human-driven behaviors. To overcome this, the paper introduces the first world modeling framework that explicitly incorporates mental variables—such as beliefs and intentions—into its core architecture. It proposes a physics–mind coupled dynamic mechanism to jointly model their interactions and integrates a modular pipeline comprising state parsing, goal observation generation, action decomposition, joint transition modeling, and branching value evaluation, leveraging large language models for interpretable mental-state reasoning. The authors also provide MENTIS, a trainable-free, testable baseline, and demonstrate through experiments on multimodal situated decision-making datasets that explicit modeling of mental states is crucial for behavior prediction, thereby revealing key bottlenecks in current approaches to mental modeling.
This work proposes a physics-informed world model that addresses the limitations of conventional video prediction approaches, which typically model dynamics directly in pixel space and struggle to capture underlying physical principles explicitly. The proposed method learns a compact, discrete "physics language" from unlabeled real-world videos through self-supervision, enabling explicit representation of world states. It adopts a "reason-then-render" paradigm: future states are predicted by performing action-conditioned sequential reasoning in this discrete latent space, followed by rendering to generate future video frames. This approach yields interpretable, physically consistent dynamics, demonstrating strong performance in both generative and perceptual tasks. Moreover, it supports interactive simulation, fine-grained action control, and zero-shot motion transfer, highlighting its capacity for structured and generalizable physical reasoning.
This work addresses the challenge of unifying prediction, planning, and irreversibility within world models. It formulates prediction as a probability measure over future trajectories and, under a local Markov assumption, employs the Onsager–Machlup action functional to decompose latent dynamics into reversible and irreversible components in path space. The authors introduce rollout-based entropy production as an operational measure of irreversibility. Through path-integral analysis, attention mechanism inspection, and small-scale model experiments, they find that attention asymmetry emerges in response to increasing data irreversibility. While symmetrization interventions suppress entropy production, they selectively impair long-horizon prediction of irreversible processes yet preserve the model’s capacity to capture relaxation dynamics.