🤖 AI Summary
This work addresses the limitations of existing deep state-space models, which rely on Gaussian assumptions and struggle to capture complex, multimodal latent dynamics. While diffusion models offer strong expressive power, they lack structured mechanisms for temporal reasoning. To bridge this gap, we propose a novel latent variable state-space model that, for the first time, integrates a non-Gaussian diffusion process into the latent state transition mechanism and enables joint training of an autoencoder and a diffusion model on sequential data. This approach overcomes the restrictive Gaussian assumption and supports unified modeling and inference of multimodal temporal dynamics. Experiments demonstrate that our model significantly outperforms state-of-the-art deep state-space models in both fitting accuracy and predictive performance on synthetic time series with complex transition characteristics.
📝 Abstract
In many domains, practitioners seek models that produce accurate forecasts while faithfully capturing latent system dynamics. Existing approaches typically sacrifice one of these goals: deep state space models often assume Gaussian latent transitions, limiting fit and forecasting, while diffusion models are highly expressive but lack principled inference for the underlying dynamics. To combine the strengths of both, we introduce the Diffusion-Driven State Space Model (DDSSM), which replaces the conventional Gaussian transition distribution with a diffusion model. Our DDSSM resolves the open problem of how to jointly train an autoencoder and a diffusion model on sequential data, thereby extending the literature on latent diffusion models for time series. Moreover, we find that the DDSSM empirically outperforms a state-of-the-art deep SSM at fitting and forecasting a simulated time series with multimodal transitions.