First Learn, Then Memorize: The Spectral Bias of Diffusion Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the transition from generalization to memorization during the late-stage training of diffusion models. Grounded in Neural Tangent Kernel (NTK) theory, this work reveals that this phenomenon stems from a dual-timescale mechanism within the spectral structure: multiple noisy replicas induce a secondary cluster of small eigenvalues in the NTK spectrum, effectively decoupling global feature directions from sample-specific noise and thereby delaying memorization. These findings are substantiated through high-dimensional limiting spectral analysis, bias-variance decomposition, and empirical validation on U-Net architectures. Furthermore, this paper proposes causal intervention strategies, including rank truncation and L2 regularization, to actively regulate this spectral dynamic. The proposed methods effectively suppress memorization and enhance generative generalization performance.
📝 Abstract
Diffusion models trained on a finite dataset first learn to generate novel, high-quality samples and only much later collapse onto their training set. We identify the mechanism behind this separation of timescales and the object that probes it. The training dynamics of the score function are governed---exactly, and at any width---by the Gram matrix of the Neural Tangent Kernel (NTK) evaluated on the noisy training data, so the timescales of generalization and of memorization must be encoded in its spectrum. We show that they are, and that the structure responsible has no analogue in standard kernel settings. The use of multiple noise realizations per sample ($m$ noised copies at a fixed noise level) in the score-matching loss is what restructures the Gram matrix spectrum into two distinct parts. The first, of large eigenvalues, carries the global features of the target distribution and is present already for $m=1$. The second, which the repeated noising creates, consists of the smallest eigenvalues and is supported on eigenvectors aligned with the sample-specific noise directions; it sets a memorization timescale parametrically larger in the training set size $n$. We establish this picture on two fronts. Analytically, we solve the spectrum in the lazy high-dimensional limit for both linear ($n \asymp d$) and polynomial ($n \asymp d^k$) sample complexities, and prove through a bias--variance decomposition that the first bulk minimizes the approximation error while the second drives the error associated with memorization. Empirically, we show the same two-bulk structure in Convolutional NTKs on CelebA and in finite-width U-Nets trained well beyond the lazy regime, and we make the link causal: truncating the Gram matrix at rank $r$ tunes the generalization--memorization transition, and an $L_2$ penalty targeting the second bulk suppresses memorization in feature-learning U-Nets.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Models
Spectral Bias
Memorization
Generalization
Neural Tangent Kernel
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion Models
Neural Tangent Kernel
Spectral Bias
Generalization-Memorization
Score Matching
🔎 Similar Papers
R
Raphaël Urfin
Laboratoire de Physique de l’École normale supérieure, ENS, Université PSL, CNRS, Sorbonne Université, Université Paris Cité, F-75005 Paris, France
T
Tony Bonnaire
Université Paris-Saclay, CNRS, Institut d’Astrophysique Spatiale, 91405 Orsay, France
Giulio Biroli
Giulio Biroli
Professor of Theoretical Physics, ENS Paris
Statistical PhysicsCondensed MatterComplex Systems
M
Marc Mézard
Department of Computing Sciences, Bocconi University, Milano, Italy