Score
Design, build, and evaluate probabilistic models that implement diffusion processes in a learned latent space and support both forward (predictive/generative) and reverse (inference/reconstruction of past or hidden) dynamics. Extend these latent-space diffusion models into autoregressive architectures that generate or score sequences of latent states (latent trajectories) and analyze their temporal consistency and conditioning behavior.
Existing diffusion model tutorials predominantly focus on Euclidean spaces, lacking a unified perspective for discrete state spaces (e.g., categorical data) and failing to systematically elucidate theoretical connections between continuous and discrete diffusion processes. Method: We propose a self-contained diffusion framework over general state spaces, unifying forward processes via stochastic differential equations (SDEs) and the Fokker–Planck equation for continuous domains, and via continuous-time Markov chains (CTMCs) and the master equation for discrete domains—both formalized through Markov kernels. We derive a unified variational objective (ELBO), explicitly characterizing how distinct noise schedules affect reverse dynamics modeling. Contribution: This work establishes, for the first time, a rigorous theoretical bridge across continuous and discrete domains. It provides a reusable derivation paradigm, proof toolkit, and pedagogical pathway, delivering a compact, general foundation for both fundamental understanding and algorithmic design of diffusion models.
Diffusion models incur substantial computational overhead, hindering their deployment in real-time physical simulation. To address this, we propose training diffusion models in the latent space of a learned autoencoder—rather than in pixel space—to drastically reduce dimensionality and accelerate inference. Experiments demonstrate that even with a 1000× latent-space compression ratio, our approach preserves high temporal fidelity in system evolution. Compared to non-generative baselines, it achieves superior predictive diversity and robustness under distributional shift. Our key contribution lies in the first systematic validation of latent-space diffusion modeling for dynamical system simulation. By jointly optimizing the encoder architecture, diffusion backbone, and training strategy, we reduce computational cost by over 90% while simultaneously improving generative quality and uncertainty quantification capability.
To address the challenge of density modeling and sampling on low-dimensional manifolds embedded in high-dimensional data, this paper proposes a dual-path framework integrating manifold learning with diffusion-based generation. First, Diffusion Maps are employed to uncover the intrinsic low-dimensional manifold structure. Second, a score-matching diffusion model is constructed—or equivalently, an Itô stochastic differential equation (SDE) is solved—directly on the learned manifold for latent-space sampling. Finally, Double Diffusion Maps—novelly applied here for unsupervised, differentiable high-dimensional reconstruction rather than dynamical dimensionality reduction—are used to lift samples back to the ambient space. The method avoids explicit manifold parameterization and enables end-to-end density modeling on unknown manifolds. Evaluated on benchmark tasks and multiscale materials datasets, it achieves significant improvements in sampling fidelity and manifold consistency, while maintaining both accuracy and computational efficiency.
This work addresses the limitations of existing deep state-space models, which rely on Gaussian assumptions and struggle to capture complex, multimodal latent dynamics. While diffusion models offer strong expressive power, they lack structured mechanisms for temporal reasoning. To bridge this gap, we propose a novel latent variable state-space model that, for the first time, integrates a non-Gaussian diffusion process into the latent state transition mechanism and enables joint training of an autoencoder and a diffusion model on sequential data. This approach overcomes the restrictive Gaussian assumption and supports unified modeling and inference of multimodal temporal dynamics. Experiments demonstrate that our model significantly outperforms state-of-the-art deep state-space models in both fitting accuracy and predictive performance on synthetic time series with complex transition characteristics.
This work addresses the challenge of conditional sampling in generative diffusion models for Bayesian inverse problems. It systematically surveys and unifies two dominant paradigms: end-to-end methods based on the joint distribution, and decoupled approaches combining a pre-trained marginal distribution with an explicit likelihood model. We propose, for the first time, a theoretically consistent unified framework that integrates Monte Carlo sampling, diffusion process reweighting, conditional probability construction, and fine-tuning techniques—rigorously characterizing the underlying assumptions and intrinsic relationships among these methods. The framework bridges theoretical gaps across disparate conditional generation strategies and delivers a scalable, interpretable, and theoretically grounded toolkit for conditional sampling in scientific computing inverse problems, including image reconstruction and physics-based simulation.
This work addresses the lack of a general framework for modeling time-varying latent states in existing generative models, which often rely on auxiliary stochastic processes that are difficult to sample. The authors propose a novel approach that treats observation generation as a deterministic mapping of a tractable Markov process, employing an image-space stochastic process generator whose one-time marginal distribution matches that of a target projected process. The key innovation lies in extending Generator Matching—previously limited to static latent variables—to time-varying latent processes for the first time. By integrating stochastic process theory, Markov projections, and flow matching techniques, the method establishes a unified generative modeling framework. This framework not only subsumes existing models with discrete latent processes as special cases but also accommodates a broader class of time-varying latent conditions while rigorously ensuring consistency between the generated and target marginal distributions.
Existing stochastic process models often struggle to perform efficient conditional sampling under nonlinear observations, non-Gaussian likelihoods, or global constraints, typically requiring bespoke and complex algorithms. This work proposes the first universal conditioning framework that requires no training and avoids neural network approximations. The approach represents stochastic processes via deterministic mappings of tractable latent innovation variables, recasting conditional sampling as an inference problem in latent space. Exact solutions are achieved through backward-time stochastic differential equations (SDEs). The method accommodates a broad class of heterogeneous processes—including spatial priors, nonlinear dynamics, stochastic partial differential equations, extreme-value processes, and discrete-state processes—and enables high-fidelity conditional sampling on a single CPU within seconds, achieving both theoretical exactness and computational efficiency.
This work addresses a critical limitation in existing diffusion-based decision-making approaches, which often neglect the evolving latent dynamics of environments, thereby struggling to accurately model state transitions, reward structures, and high-level behaviors. To overcome this, the paper proposes the first unified causal diffusion framework with theoretical identifiability guarantees, capable of explicitly inferring and jointly learning latent dynamics and observation interactions from limited data. By integrating a modular architecture with a novel identification technique based on short temporal segments, the method adaptively captures changes in dynamics, rewards, and latent actions. Empirical evaluations on simulated control and robotic benchmark tasks demonstrate substantial improvements in both latent inference accuracy and policy generalization.
This work addresses the challenge of generating conditional stochastic processes in continuous spatiotemporal settings from arbitrary observation subsets—such as irregularly sampled data or future frames—by proposing an autoregressive generative framework based on non-Markovian diffusion bridges. The method unifies physical time and state evolution within a single continuous stochastic differential equation (SDE), innovatively initializing from neighboring states, injecting noise proportionally to temporal intervals, and explicitly embedding time into the SDE dynamics. The SDE is derived via path-space measure transformation, and training is performed using a path- and time-dependent denoising score matching algorithm. Empirical evaluations on video generation and weather forecasting demonstrate significant improvements over existing approaches, particularly under low-step sampling and irregular conditioning scenarios.
This study addresses the challenge of decoding latent dynamic structures underlying large-scale neuronal population activity by proposing a unified latent variable modeling framework that, for the first time, jointly integrates three core tasks: single-region dynamics modeling, inter-regional communication analysis, and behavioral alignment. The approach combines classical state-space models with cutting-edge deep generative architectures—including Transformers, diffusion models, and neural ordinary differential equations—to systematically construct a taxonomy and establish clear evaluation benchmarks. Emphasizing critical challenges such as causal inference and directional connectivity, this work provides both theoretical foundations and methodological tools for interpretable brain dynamics analysis and robust neural decoding.