Identifiable World Models from Pretrained Diffusion Representations

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation that diffusion world models, despite accurate predictions, struggle to recover identifiable latent state variables and causal interactions. To this end, we propose a contrastive diffusion alignment framework that establishes, for the first time, the transferability of auxiliary-variable nonlinear independent component analysis theory to frozen pretrained diffusion latent spaces. Under temporal or generalized contrastive learning assumptions, a lightweight mapping is learned via transition Jacobian sparsity constraints, enabling the extraction of latent dynamical states without retraining the generative backbone. Empirically, the proposed approach achieves near-perfect block-wise state recovery and lag-graph reconstruction in physical and robotic video systems, while precisely characterizing intervention-related dynamical responses. These results demonstrate that structured, interpretable state coordinates can be effectively endowed to diffusion models.
📝 Abstract
Diffusion-based world models can generate and predict trajectories in high-dimensional dynamical systems, but predictive accuracy does not imply that their latent coordinates recover the underlying state variables or causal interactions. We ask whether a frozen pretrained diffusion model can be equipped with identifiable coordinates without retraining its generative backbone. We show that auxiliary-variable nonlinear ICA guarantees can be transferred to Contrastive Diffusion Alignment (ConDA), which learns only a lightweight alignment map on top of frozen diffusion latents. Under standard TCL/GCL assumptions, the aligned representation identifies latent dynamical states up to permutation and componentwise invertible transformations, preserves the latent dynamic structural causal model, and reduces lagged graph recovery to transition-Jacobian sparsity. We evaluate TCL-, GCL-, and CEBRA-based ConDA against TDRL, CaRiNG, IDOL, temporal SuaVE, and iVAE across physical and robotic video systems. TCL and GCL achieve near-perfect blockwise state recovery and competitive lagged graph recovery, including exact recovery in a simulated falling-body system. In a simulated bipedal robot, learned dynamics recover the sign and temporal structure of responses to held-out control perturbations. These results show that a frozen generative diffusion model can be equipped with coordinates that are identifiable, structurally interpretable, and useful for analyzing intervention-relevant dynamics.
Problem

Research questions and friction points this paper is trying to address.

World Models
Identifiability
Diffusion Models
Nonlinear ICA
Causal Discovery
Innovation

Methods, ideas, or system contributions that make the work stand out.

Contrastive Diffusion Alignment
Identifiable World Models
Nonlinear ICA
Pretrained Diffusion Models
Causal Discovery
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Ruchi Sandilya
Department of Psychiatry and Brain & Mind Research Institute, Weill Cornell Medicine, New York, NY, USA
C
Conor Liston
Department of Psychiatry and Brain & Mind Research Institute, Weill Cornell Medicine, New York, NY, USA
Logan Grosenick
Logan Grosenick
Assistant Professor (tenure track) at Cornell University
Machine Learning/AIPsychiatryNeuroimagingMultiomics