🤖 AI Summary
This study addresses the long-standing lack of theoretical understanding regarding memorization and generalization mechanisms in conditional diffusion models. Based on a high-dimensional random feature conditional score model, this work employs random feature theory and proportional limit analysis to derive asymptotic expressions for training and test losses, which are further validated through experiments with U-Net architectures. Notably, it introduces the concept of “malignant generalization” for the first time, revealing the intrinsic mechanism whereby increasing network width under overparameterization improves conditional mean prediction yet degrades variance estimation, while elucidating the critical influence of conditioning information on memorization. Ultimately, this project establishes a unified theoretical framework for conditional diffusion models and corroborates these theoretical findings on real-world data.
📝 Abstract
Conditional diffusion models generate diverse, novel, and high-quality samples under prescribed conditions. However, theoretical understanding of their memorization and generalization remains limited, while recent works have characterized these behaviors primarily in unconditional settings. In this work, we analyze a random-feature conditional score model in the high-dimensional proportional limit, deriving asymptotic expressions for training and test losses. By decomposing the test loss, we show that in the overparameterized regime, increasing model width improves prediction of the condition-dependent mean while reducing within-condition prediction variance, a phenomenon we term "malign generalization." Furthermore, analyzing the training loss reveals that more informative conditions lead to memorization of training samples at smaller widths. These theoretical findings are supported by experiments with U-Net architectures on realistic data.