🤖 AI Summary
This study addresses the limitation that diffusion world models, despite accurate predictions, struggle to recover identifiable latent state variables and causal interactions. To this end, we propose a contrastive diffusion alignment framework that establishes, for the first time, the transferability of auxiliary-variable nonlinear independent component analysis theory to frozen pretrained diffusion latent spaces. Under temporal or generalized contrastive learning assumptions, a lightweight mapping is learned via transition Jacobian sparsity constraints, enabling the extraction of latent dynamical states without retraining the generative backbone. Empirically, the proposed approach achieves near-perfect block-wise state recovery and lag-graph reconstruction in physical and robotic video systems, while precisely characterizing intervention-related dynamical responses. These results demonstrate that structured, interpretable state coordinates can be effectively endowed to diffusion models.
📝 Abstract
Diffusion-based world models can generate and predict trajectories in high-dimensional dynamical systems, but predictive accuracy does not imply that their latent coordinates recover the underlying state variables or causal interactions. We ask whether a frozen pretrained diffusion model can be equipped with identifiable coordinates without retraining its generative backbone. We show that auxiliary-variable nonlinear ICA guarantees can be transferred to Contrastive Diffusion Alignment (ConDA), which learns only a lightweight alignment map on top of frozen diffusion latents. Under standard TCL/GCL assumptions, the aligned representation identifies latent dynamical states up to permutation and componentwise invertible transformations, preserves the latent dynamic structural causal model, and reduces lagged graph recovery to transition-Jacobian sparsity. We evaluate TCL-, GCL-, and CEBRA-based ConDA against TDRL, CaRiNG, IDOL, temporal SuaVE, and iVAE across physical and robotic video systems. TCL and GCL achieve near-perfect blockwise state recovery and competitive lagged graph recovery, including exact recovery in a simulated falling-body system. In a simulated bipedal robot, learned dynamics recover the sign and temporal structure of responses to held-out control perturbations. These results show that a frozen generative diffusion model can be equipped with coordinates that are identifiable, structurally interpretable, and useful for analyzing intervention-relevant dynamics.