🤖 AI Summary
To address the limited sample quality and diversity of score-matching generative models for structured distribution modeling, this paper proposes a novel framework integrating nonlinear forward dynamics with score matching in the VAE latent space. Our key contributions are: (1) replacing conventional linear stochastic differential equations (SDEs) with nonlinear denoising score matching in the VAE latent space; (2) reformulating the cross-entropy regularization term to mitigate gradient variance explosion under small step sizes; and (3) adopting the Euler–Maruyama scheme to approximate Gaussian transitions, balancing accuracy and computational efficiency. Experiments on MNIST variants demonstrate that our method significantly improves generation speed, Fréchet Inception Distance (FID), and diversity metrics—outperforming standard latent-space generative models. This work establishes a new paradigm for efficient, high-fidelity modeling of structured data.
📝 Abstract
We present latent nonlinear denoising score matching (LNDSM), a novel training objective for score-based generative models that integrates nonlinear forward dynamics with the VAE-based latent SGM framework. This combination is achieved by reformulating the cross-entropy term using the approximate Gaussian transition induced by the Euler-Maruyama scheme. To ensure numerical stability, we identify and remove two zero-mean but variance exploding terms arising from small time steps. Experiments on variants of the MNIST dataset demonstrate that the proposed method achieves faster synthesis and enhanced learning of inherently structured distributions. Compared to benchmark structure-agnostic latent SGMs, LNDSM consistently attains superior sample quality and variability.