Improved Distributional Diffusion Models

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive multi-particle training overhead and the limited sampling flexibility imposed by global fixed scoring rules when extending Diffusion Distribution Models (DDMs) to modern image generation. We propose an efficient single-stage class-conditional generation method based on a DiT latent space architecture. By deferring particle computations to deeper Transformer layers, the approach eliminates linear overhead, while a time-dependent dynamic scoring rule schedule is designed to overcome traditional performance bottlenecks. The method supports training from scratch without requiring teacher models or self-distillation. On ImageNet-256, it achieves an FID of 4.48 with 4 steps and 2.38 with 50 steps, notably without FID degradation as sampling steps increase. Furthermore, the proposed framework successfully generalizes to text-to-image generation tasks.
📝 Abstract
Distributional Diffusion Models (DDMs) replace the standard mean-prediction denoiser with a \emph{distributional} denoiser trained via a scoring rule objective, learning a stochastic approximation to $p(x_1 \mid x_t)$ rather than its conditional mean. However, scaling DDMs to modern image-generation settings faces two obstacles: (i) multi-particle training incurs overhead that scales with the number of particles, (ii) DDMs use globally fixed scoring rule hyperparameters, forcing a single trade-off across sampling budgets. We mitigate these limitations by deferring particle expansion to late transformer layers, and the hyperparameter trade-off by introducing time-dependent scoring rule schedules informed by the dynamical regimes of~\citet{Biroli2024}. Combined with a DiT-based latent setup, these changes make DDM training practical on class-conditional ImageNet-$256^2$, achieving 4.48 FID at 4 steps and 2.38 at 50 steps with DiT-XL/2, from a single model trained from scratch in one stage, without a teacher, self-distillation or JVPs. The result is a stochastic few-step generator whose FID does not degrade as the sampling budget grows from 4 to 50 NFE, and the same recipe transfers to text-to-image generation. Code and pre-trained models available at https://github.com/CompVis/iDDM.
Problem

Research questions and friction points this paper is trying to address.

Distributional Diffusion Models
multi-particle training overhead
scoring rule hyperparameters
image generation
sampling budget
Innovation

Methods, ideas, or system contributions that make the work stand out.

Distributional Diffusion Models
Scoring Rule Schedules
Late Particle Expansion
Few-step Generation
DiT
🔎 Similar Papers
No similar papers found.