Double Descent and Malign Overfitting in Diffusion Models

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究解决了扩散模型中的恶性过拟合问题,通过实验和理论分析发现,在训练样本数固定的情况下,随着参数增加,测试误差提前增大。
📝 Abstract
Conventional wisdom in deep learning holds that overparameterization---having more parameters $p$ than training samples $n$---is benign: larger models generalize better and, even without regularization, interpolating models generalize well, the test error following a double-descent curve. One might expect the same benign overfitting for diffusion models, whose training reduces to regression, i.e. to minimizing a quadratic score-matching loss. Yet the opposite is observed: overfitting here is catastrophic, driving the model into a memorization regime. We resolve this paradox by combining experiments on U-Nets trained on CelebA with a random-features model for which we derive closed-form learning curves. We show that with a fixed number $m$ of noise realizations per training sample, an interpolation peak does occur, but at $p\sim nm$ rather than at $p\sim n$ as in standard regression. The rise of the test loss, however, sets in much earlier, at $p\sim n$, independently of $m$. This overfitting is malign because, although the implicit regularization of training is fully at work, it drives the model toward the empirical score, which memorizes the training set, rather than toward the true score. A bias-variance decomposition pinpoints the mechanism: the bias of the score estimator starts to grow at $p\sim n$; past the peak the variance decays, as in regression, whereas the bias keeps growing and both saturate at a large value. Since diffusion models are trained with $m\gg1$, the peak is pushed to very large model sizes, and therefore sit on the rising branch that precedes it, where malign overfitting is already in play. Nevertheless, overparameterization remains beneficial when paired with regularization: in the random-features theory and in U-Net experiments, optimally regularized large models---via a ridge penalty or early stopping, respectively---outperform any unregularized models.
Problem

Research questions and friction points this paper is trying to address.

overparameterization
malign overfitting
diffusion models
generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

malign overfitting
diffusion models
interpolation peak
implicit regularization
bias-variance decomposition
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Raphaël Urfin
Laboratoire de Physique de l’École normale supérieure, ENS, Université PSL, CNRS, Sorbonne Université, Université Paris Cité, F-75005 Paris, France
T
Tony Bonnaire
Université Paris-Saclay, CNRS, Institut d’Astrophysique Spatiale, 91405 Orsay, France
Giulio Biroli
Giulio Biroli
Professor of Theoretical Physics, ENS Paris
Statistical PhysicsCondensed MatterComplex Systems
M
Marc Mézard
Department of Computing Sciences, Bocconi University, Milano, Italy