Smoothing the Score Function for Generalization in Diffusion Models: An Optimization-based Explanation Framework

📅 2026-01-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Diffusion models are prone to reproducing training samples during generation due to memorization, which compromises their generalization capability. This work theoretically shows that the empirical score function consists of Gaussian scores weighted by a sharp softmax, causing individual training samples to dominate the generation process. To address this, the authors propose two novel techniques: Noise Unconditioning and Temperature Smoothing. The former adaptively adjusts sample weights, while the latter explicitly controls the softmax temperature; together, they yield a smoothed approximation of the score function, ensuring that sampling is guided by the local data manifold rather than isolated points. Experimental results validate the theoretical analysis, demonstrating that the proposed methods significantly improve generalization across multiple datasets while preserving high-quality generation.

Technology Category

Computer Vision: Diffusion Models for VisionMachine Learning: Deep Generative Models & AutoencodersNatural Language Processing: Generation

Application Category

User Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Diffusion models achieve remarkable generation quality, yet face a fundamental challenge known as memorization, where generated samples can replicate training samples exactly. We develop a theoretical framework to explain this phenomenon by showing that the empirical score function (the score function corresponding to the empirical distribution) is a weighted sum of the score functions of Gaussian distributions, in which the weights are sharp softmax functions. This structure causes individual training samples to dominate the score function, resulting in sampling collapse. In practice, approximating the empirical score function with a neural network can partially alleviate this issue and improve generalization. Our theoretical framework explains why: In training, the neural network learns a smoother approximation of the weighted sum, allowing the sampling process to be influenced by local manifolds rather than single points. Leveraging this insight, we propose two novel methods to further enhance generalization: (1) Noise Unconditioning enables each training sample to adaptively determine its score function weight to increase the effect of more training samples, thereby preventing single-point dominance and mitigating collapse. (2) Temperature Smoothing introduces an explicit parameter to control the smoothness. By increasing the temperature in the softmax weights, we naturally reduce the dominance of any single training sample and mitigate memorization. Experiments across multiple datasets validate our theoretical analysis and demonstrate the effectiveness of the proposed methods in improving generalization while maintaining high generation quality.
Problem

Research questions and friction points this paper is trying to address.

memorization
diffusion models
generalization
score function
sampling collapse
Innovation

Methods, ideas, or system contributions that make the work stand out.

score function smoothing
memorization mitigation
temperature smoothing
noise unconditioning
diffusion model generalization
X
Xinyu Zhou
Department of Computer Sciences, University of Wisconsin Madison, WI, USA
J
Jiawei Zhang
Department of Computer Sciences, University of Wisconsin Madison, WI, USA
S
Stephen J. Wright
Department of Computer Sciences, University of Wisconsin Madison, WI, USA