Score
Designs and implements fast samplers for diffusion-based generative models that reduce the number of denoising iterations while preserving output fidelity. This includes developing and analyzing accelerated update rules (e.g., DDIM and diffusion-posterior schemes), measurement- or process-constrained sampling methods, and regularization or roughness-constrained strategies to integrate conditioning and enforce priors during few-step generation.
This paper investigates how diffusion models automatically exploit unknown low-dimensional structures inherent in target distributions to accelerate sampling. Focusing on DDIM and DDPM—two canonical samplers—we analyze their convergence rates under exact score estimation via total variation distance. Method: We combine probabilistic distance analysis, iterative complexity theory, and nonparametric distributional assumptions—without requiring smoothness or log-concavity. Contribution/Results: We provide the first rigorous proof that DDIM-style samplers are adaptive to unknown low-dimensional manifolds. We establish tight upper and lower bounds, revealing that the sampling coefficients proposed by Ho et al. (2020) and Song et al. (2020) are (nearly) necessary for such adaptivity. Our analysis shows that both DDIM and DDPM achieve an iteration complexity of $O(k/varepsilon)$ (up to logarithmic factors), significantly improving upon prior total-variation convergence bounds for DDPM and extending applicability to broad classes of unstructured target distributions.
Diffusion probabilistic models (DPMs) suffer from slow sampling and numerical divergence under high classifier-free guidance (CFG) scales. To address this, we propose a stable and efficient high-order ODE-guided sampling framework. Our method introduces three key innovations: (1) the first high-order explicit solver specifically designed for data-prediction DPMs; (2) a threshold-based distribution correction mechanism to suppress gradient explosion at large CFG scales; and (3) a multi-step adaptive step-size reduction strategy to ensure numerical stability. Evaluated on both pixel-space and latent-space DPMs, our approach generates high-fidelity samples in only 15–20 function evaluations—over 5× faster than DDIM—while maintaining robust convergence across CFG=15–25. The implementation is publicly available.
Diffusion models suffer from high inference costs, scaling linearly—O(d) or O(T)—with data dimensionality or time steps, hindering practical deployment. To address this, we propose a block-parallel Picard iteration framework for accelerated sampling. We provide the first rigorous proof that this approach achieves sublinear time complexity—specifically, ~O(poly log d)—thereby breaking the fundamental linear bottleneck. Our method integrates theoretical analysis grounded in the generalized Girsanov theorem with a dual-path design compatible with both stochastic differential equations (SDEs) and probability flow ordinary differential equations (ODEs). This enables scalable, efficient parallelization across large GPU clusters. Experiments on scientific modeling and image generation tasks demonstrate substantial speedups in sampling efficiency—up to orders of magnitude—without compromising sample quality. The framework establishes a new paradigm for high-dimensional diffusion sampling, uniquely combining theoretical guarantees with engineering feasibility.
To address the high computational cost of diffusion model sampling—specifically the large number of neural function evaluations (NFEs) required—this paper proposes LD3, a lightweight, plug-and-play framework. LD3 is the first method to formulate timestep selection as a learnable, parameterized process without modifying or retraining the backbone diffusion model, and it is compatible with mainstream solvers such as DDIM and Heun. Leveraging gradient-driven timestep optimization and adaptive ODE discretization, LD3 achieves significantly improved sampling efficiency while maintaining theoretical convergence guarantees. Trained in only 5–10 minutes on CIFAR-10 and AFHQv2, LD3 attains FID scores of 2.38 and 2.27 using merely 10 NFEs—substantially reducing computational cost and improving generation quality over baselines. Its core innovation lies in decoupling the learning of time discretization from model weight updates, enabling efficient, robust, and architecture-agnostic acceleration.
This work addresses the convergence rate of score-based diffusion models (e.g., DDPM) under weak assumptions. Methodologically, it formulates the problem via stochastic differential equations (SDEs) and introduces a refined error propagation analysis, tight control of score estimation errors, and intrinsic-dimension-aware coefficient design. Theoretically, it establishes the first rigorous convergence guarantee—without requiring smoothness, strong convexity, or prior knowledge of the data distribution—showing that total variation distance converges at rate $O(d/T)$ given only $ell_2$-accurate score estimates, which is dimensionally optimal for any target distribution with finite first moment. Moreover, the analysis reveals that DDPM automatically adapts to unknown low-dimensional structure, achieving an improved rate of $O(k/T)$, where $k$ denotes the intrinsic data dimension. This is the first result providing dimensionally explicit, minimally assumed, and provably tight convergence bounds, thereby offering foundational theoretical support for the efficiency and generalization capability of diffusion models.
Existing theoretical analyses of diffusion generative models are fragmented and suffer from inconsistent notation, hindering unified understanding and principled development. Method: This work establishes a rigorous, unified mathematical framework grounded in fundamental properties of Gaussian distributions. It systematically derives the closed-form marginal distribution of the forward noising process, the analytical form of the reverse posterior, and the variational lower bound, ultimately yielding an optimization objective equivalent to noise prediction. Contribution/Results: The framework reveals the intrinsic equivalence between DDIM and rectified flow; provides a unified probabilistic interpretation of classifier-guided and classifier-free guidance; and integrates SDE/ODE formulations, the Fokker–Planck equation, flow matching, and multi-scale modeling—ensuring both theoretical coherence and practical implementability. Validated on mainstream models including Stable Diffusion, the framework enables efficient sampling and precise modeling while unifying disparate theoretical perspectives.
Diffusion models suffer from degraded image quality when constrained to a limited number of reverse denoising steps (e.g., 4–8). To address this, we propose a plug-and-play, model-agnostic parallel dual-sampler framework. Our method runs two independent samplers in parallel at adjacent timesteps and employs a latent-space adaptive fusion strategy to jointly optimize the latent variables—avoiding performance degradation caused by naive averaging. Experiments demonstrate that, with only marginal computational overhead, our approach significantly improves few-step generation quality across diverse diffusion models—including Stable Diffusion and DDPM. Gains are consistently validated by both automated metrics (FID, LPIPS) and human evaluation. The core contribution lies in revealing that “fewer but smarter” parallel sampling outperforms simply increasing sampler count, thereby establishing a lightweight, general-purpose paradigm for high-fidelity few-step generation.
This work addresses the susceptibility of existing stochastic differential equation (SDE)-based guided diffusion methods to estimation errors during posterior sampling, which often leads to gradient bias and inconsistent generation. To mitigate these issues, the authors propose a plug-and-play backward denoising strategy that incorporates an additional denoising step combined with adaptive Monte Carlo sampling (ABMS). This approach effectively reduces cross-condition interference, thereby enhancing both the accuracy and stability of guidance. Notably, the method is compatible with high-order diffusion samplers and requires no modifications to the underlying model architecture. Experimental results across diverse tasks—including handwritten trajectory generation, image inpainting, and molecular inverse design—demonstrate consistent and significant improvements in generation quality over current state-of-the-art approaches.
Existing diffusion models often suffer from low constraint satisfaction rates or degraded sample quality when generating samples under complex feasibility constraints, primarily due to distributional mismatches between training and sampling phases. This work proposes a trajectory-aware fine-tuning framework based on online rollouts, which integrates constraint guidance during training and, for the first time, incorporates a numerical integration perspective into the diffusion process. By end-to-end differentiating denoising trajectories under a fixed noise schedule, the method explicitly exposes constraint violations, thereby aligning the training and sampling distributions. Combining constraint-aware guidance, differentiable noise scheduling, and an online rollout mechanism, the approach significantly improves constraint satisfaction across multiple tasks while maintaining generation quality on par with current state-of-the-art methods.
Existing DDPM sampling relies on exponential Euler discretization, suffering from convergence complexity at least O(d) or O(√I₀), severely limiting scalability in high dimensions. Method: We propose DDRaM—a novel SDE integration method for DDPMs based on stochastic midpoints—introducing an auxiliary random midpoint to enhance discretization accuracy while preserving the standard DDPM sampling mechanism. Under Lipschitz smoothness of the score function, DDRaM achieves improved numerical fidelity without altering the underlying generative process. Contribution/Results: We establish, for the first time, a sublinear iteration complexity bound for vanilla DDPM sampling: only Õ(√d) score evaluations suffice to attain ε-accuracy. This breaks prior limitations requiring ODE approximations or sampler modifications, better aligning with practical deployment. Our analysis integrates the “displacement composition rule” framework with log-concave sampling techniques. Empirical validation on pre-trained image generation models confirms substantial speedup.