Score
Analyze and prove properties of SDE-based diffusion samplers: derive stochastic differential identities and martingale relations, establish existence and uniqueness of invariant measures, and prove ergodicity and quantitative convergence rates. Analyze and bound discretization, sampling, and approximation error when mapping continuous-time sampling dynamics to practical discrete algorithms, and produce proof roadmaps for key theorems about sampler convergence.
A theoretical-practical gap persists in score-based diffusion models. Method: We propose a unified, reproducible SDE-based modeling framework that systematically integrates score matching, SDE/ODE solvers, denoising score estimation, and consistency modeling; notably, we introduce reinforcement learning into diffusion sampling for inference-path optimization. Contributions: (1) We establish theoretical consistency between sampling and score estimation under the SDE formulation; (2) we provide concise proofs of key theorems alongside practical algorithmic implementation guidelines; (3) we release modular, open-source code enabling rapid validation and extension to novel architectures. This work bridges the efficiency of score matching with scalable, RL-enhanced inference, delivering a foundational toolkit that balances theoretical rigor and engineering practicality for the design, analysis, and application of diffusion models.
This paper investigates the dynamical behavior of Thompson sampling under the joint asymptotic regime of small gaps—where arm mean differences scale as $O(sqrt{gamma})$—and long horizons—where the time horizon scales as $O(1/gamma)$. Using weak convergence analysis, we derive, for the first time from first principles, a diffusion approximation: we prove that the discrete-time update process converges weakly to an explicit stochastic differential equation (SDE) and its associated random ordinary differential equation (ODE). This limiting characterization unifies the asymptotics of diverse Thompson sampling variants—including those based on exponential families and bootstrap resampling—and reveals intrinsic robustness under model misspecification. Our results establish a universal limit theory for Thompson sampling in the small-gap regime and provide a novel analytical framework for continuous-time modeling and robustness analysis of bandit algorithms.
This work unifies diffusion sampling and stochastic localization under a single theoretical framework, addressing the lack of rigorous theoretical connections between them and the limited applicability of existing algorithms. Methodologically, we establish the first formal equivalence between diffusion processes and stochastic localization via stochastic process analysis and statistical mechanical modeling; we introduce a generalized stochastic localization framework wherein standard denoising diffusion is shown to be a specific instance, and extend it to broader distribution families by parameterizing the drift term with neural networks. Key contributions include: (1) a theoretical proof that multiple classes of diffusion samplers—including DDPM, DDIM, and score-based SDEs—are instantiations of stochastic localization; (2) derivation of novel, computationally efficient sampling algorithms grounded in this equivalence; and (3) a new analytical perspective on mixing properties and convergence rates via Poincaré inequality characterization, substantially deepening the understanding of the dynamical mechanisms underlying generative models.
This work addresses discrete sampling and linear functional estimation on compact Riemannian manifolds via intrinsic Langevin diffusion. For target distributions μ_φ ∝ e⁻ᵠ with only C¹-smooth potential φ—without requiring second-order smoothness or curvature bounds—we propose a retraction-based discretization of the Langevin diffusion as a Markov chain. Our contribution is threefold: (i) we establish first-order optimal step-size-dependent bounds on bias and mean-squared error without assuming φ ∈ C² or bounded curvature; (ii) we prove that the stationary measure of the discrete chain converges to μ_φ at rate O(h) in Wasserstein distance, where h is the step size; (iii) the analysis applies even when the exponential map lacks a closed-form expression and extends to non-compact manifolds. Numerical experiments confirm the algorithm’s efficacy on both positively and negatively curved manifolds for log-concave and related target distributions.
This work addresses the theoretical behavior and numerical error control of diffusion models under Gaussian data distributions. We systematically analyze the analytical solutions of the backward stochastic differential equation (SDE) and probability flow ordinary differential equation (ODE), and— for the first time—establish a rigorous, term-wise decomposition and exact quantification framework for four fundamental error sources: initialization, truncation, discretization, and score approximation, all measured in Wasserstein distance. Leveraging properties of Gaussian processes and SDE/ODE theory, we prove that all analytical solutions and mainstream discretization schemes remain Gaussian processes, enabling closed-form error computation directly in the data space. This yields the first complete error spectrum for diffusion sampling under Gaussian assumptions, eliminating reliance on proxy metrics (e.g., Inception Score) and permitting direct verification of sampler optimality. Our results provide a strict, computationally tractable theoretical benchmark for both analysis and algorithm design of diffusion models.
This work investigates the long-term approximation accuracy of stochastic gradient Langevin dynamics (SGLD) to continuous Langevin diffusion, focusing on uniform-in-time error bounds for the Kullback–Leibler (KL) divergence and Wasserstein/total variation distances between their invariant measures. Leveraging a synthesis of stochastic differential equation analysis, information-theoretic entropy estimation, and diffusion approximation theory under non-convex potentials, we establish, for the first time, a sharp, uniform-in-time $O(eta^2)$ upper bound on the KL divergence for step size $eta$. This directly implies $O(eta)$ bounds on the Wasserstein and total variation distances between the invariant measures. The results hold for general non-convex potentials and accommodate variable step sizes—significantly improving upon prior $O(eta)$ KL bounds. To date, this provides the strongest theoretical guarantee for the stability and statistical fidelity of SGLD in Bayesian inference and sampling.
This work addresses the challenge faced by beginning graduate students who lack prior exposure to stochastic differential equations and diffusion models by proposing a hierarchical pedagogical framework that systematically constructs the mathematical foundations of diffusion models. Starting from a sampling perspective, it integrates core definitions, key estimates under simplified assumptions, and proof strategies for cutting-edge theorems, thereby bridging classical sampling dynamics with modern diffusion samplers. The material synthesizes probability theory, stochastic differential equations, stochastic numerical methods, and diffusion process theory into a self-contained, proof-oriented curriculum. This approach maintains mathematical rigor while significantly enhancing accessibility, enabling students without prerequisite knowledge to grasp the sampling mechanisms, error analysis, and inference control principles underlying diffusion models.
This work establishes a unified theoretical framework for diffusion models from the perspective of differential equations. Starting from a conditional Gaussian forward process, it derives the corresponding forward stochastic differential equation (SDE) and ordinary differential equation (ODE), and constructs a dynamical system that transports the data distribution to a standard Gaussian prior via marginalization. The framework then introduces a reverse SDE and a probability flow ODE, both driven by the marginal score function, thereby unifying score matching and noise prediction objectives. It rigorously demonstrates the equivalence of DDPM and DDIM in their training objectives while clarifying their fundamental distinction in sampling mechanisms—DDPM corresponds to a discretized reverse SDE, whereas DDIM implements a reverse ODE. Furthermore, the framework seamlessly incorporates mainstream sampling techniques such as DPM-Solver and classifier guidance, providing a coherent and rigorous continuous-time foundation for diffusion models.
This work addresses the lack of a clear understanding regarding the direct discretization link between the Föllmer process and denoising diffusion probabilistic model (DDPM) samplers. By interpreting the Föllmer process as a time-compressed, augmented form of the DDPM reverse stochastic differential equation (SDE), this study establishes—for the first time—a systematic correspondence between the two at the discretization level. Building on this perspective, we develop a novel theoretical framework for analyzing sampling errors in DDPMs, which naturally yields optimal hyperparameter configurations. Furthermore, our approach leads to a modest yet meaningful improvement over the current best-known error bounds, achieved through a more streamlined derivation.
This study addresses the lack of first-order theoretical foundations and convergence guarantees in diffusion model sampling. By integrating stochastic differential equations (SDEs), Langevin dynamics, and non-convex optimization theory, it establishes a first-order analytical framework for diffusion models. The work reveals the contraction advantages of SDEs over ordinary differential equations (ODEs) and introduces a local score consistency certificate that does not require global convexity. Specifically, it proves that the reverse SDE flow exhibits exponential contraction in Fisher divergence under strongly convex potentials. Furthermore, it derives first-order stationarity bounds following discretization, yielding explicit exponential convergence rates and sampling convergence guarantees at the level of average gradient norms.
This study addresses the challenge of estimating diffusion parameters in stochastic differential equation (SDE) models when data and model are compatible only at specific scales. The authors propose an adaptive subsampling method based on the statistics of monotonic runs. By demonstrating that, for a broad class of additive-noise SDEs, the length of monotonic runs at infinitesimal scales approximately follows a geometric distribution with success probability 1/2, they establish a general criterion for selecting the subsampling rate without relying on multiscale diffusion asymptotics. The optimal sampling scale matching the SDE’s infinitesimal behavior is automatically determined solely from the statistical properties of monotonic increasing or decreasing segments in the observed time series. Validation on surrogate modeling of fiber lay-down trajectories in nonwoven fabric production demonstrates that the method yields highly accurate and model-consistent diffusion parameter estimates, proving effective in real-world industrial applications.