🤖 AI Summary
Schrödinger Bridge (SB) methods in diffusion modeling require estimating intractable forward score functions and rely on costly implicit trajectory simulations, severely limiting scalability. To address this, we propose the Variational Schrödinger Diffusion Model (VSDM), the first SB framework integrating variational inference: it linearizes the forward score to enable simulation-free backward score training. We theoretically establish convergence of the variational score to the true backward score along sampling trajectories, eliminating the need for warm-up initialization and enhancing hyperparameter robustness. Experiments demonstrate that VSDM yields straighter, more anisotropic shape generation trajectories; achieves state-of-the-art performance on unconditional CIFAR-10 generation; effectively handles conditional time-series modeling; and significantly improves training efficiency and scalability in large-scale settings.
📝 Abstract
Schr""odinger bridge (SB) has emerged as the go-to method for optimizing transportation plans in diffusion models. However, SB requires estimating the intractable forward score functions, inevitably resulting in the costly implicit training loss based on simulated trajectories. To improve the scalability while preserving efficient transportation plans, we leverage variational inference to linearize the forward score functions (variational scores) of SB and restore simulation-free properties in training backward scores. We propose the variational Schr""odinger diffusion model (VSDM), where the forward process is a multivariate diffusion and the variational scores are adaptively optimized for efficient transport. Theoretically, we use stochastic approximation to prove the convergence of the variational scores and show the convergence of the adaptively generated samples based on the optimal variational scores. Empirically, we test the algorithm in simulated examples and observe that VSDM is efficient in generations of anisotropic shapes and yields straighter sample trajectories compared to the single-variate diffusion. We also verify the scalability of the algorithm in real-world data and achieve competitive unconditional generation performance in CIFAR10 and conditional generation in time series modeling. Notably, VSDM no longer depends on warm-up initializations and has become tuning-friendly in training large-scale experiments.