Early Signatures of Memorization in Diffusion Models via Basin Geometry and Cyclic Denoising

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of detecting memorization in diffusion models prior to generation by proposing an early auditing method grounded in energy landscape geometry. The research reveals that memorization manifests before output generation, introducing the concept of "latent memorization." By identifying degenerate attractors, it refines the conventional logic that equates memorization solely with basin residence time. Through score divergence, basin volume analysis, and cyclic denoising techniques to probe local basins surrounding training samples, the method enables early identification and isolation of memorized content. Supported by theoretical proofs and multi-dataset experiments, this approach successfully recovers training images on CelebA without explicit replication and achieves an AUC of 0.944 on Stable Diffusion, significantly advancing the detection window for model memorization.
📝 Abstract
Diffusion models generalize early in training and later reproduce individual training samples. Standard tests detect memorization only once one-shot generation produces near-copies, leaving a released model unaudited until its outputs fail. We show that memorization is encoded in the geometry of the learned energy landscape before it appears in generated samples, a state we call latent memorization. Using score divergence and basin volume, we find that localized basins form around training samples and separate them from held-out samples before the first memorized sample appears, with an onset that follows the same $O(n)$ scaling as the memorization time. We probe these basins with cyclic denoising, which repeatedly applies partial noising and denoising. Under the exact empirical score, we prove that cycling started near an isolated training sample recovers it and returns to it over any finite number of cycles with high probability. In trained models, cycling recovers training images from CelebA and CIFAR-10 checkpoints whose one-shot samples contain no copies, and at a CelebA checkpoint with 0.1% one-shot copies, 500 cycles raise the memorized fraction above 30%. Cycling also reveals degenerate attractors that match no single training image and fade as training proceeds, so residence in a basin does not by itself imply memorization. These findings hold on a Gaussian mixture, CelebA, and CIFAR-10 across optimizers, architectures, noise schedules, and training-set sizes, and extend to off-the-shelf Stable Diffusion v1.4, where the cycled conditional-unconditional divergence gap separates memorized from non-memorized prompts with an AUC of 0.944 and a TPR of 0.866 at 1% FPR. More broadly, what a diffusion model has memorized is a property of the geometry and stability of its learned distribution, and assessing it requires examining this structure rather than generated outputs alone.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Models
Memorization
Early Detection
Energy Landscape
Latent Memorization
Innovation

Methods, ideas, or system contributions that make the work stand out.

latent memorization
cyclic denoising
basin geometry
score divergence
diffusion models
🔎 Similar Papers
No similar papers found.