🤖 AI Summary
This study addresses the high false positive rates and unreliable detection in existing generative model memorization auditing, which stem from the absence of well-defined null distribution assumptions. To overcome this limitation, this work formulates null hypotheses at both holistic and single-image levels and proposes a permutation testing framework grounded in sample exchangeability. A matched-control calibration strategy is further introduced to accurately estimate the null distribution. Additionally, rigorous multiple hypothesis testing is achieved by integrating false discovery rate (FDR) control with intersection Euler characteristic profile statistics. The proposed approach effectively rectifies the inflated false positive issue inherent in current benchmarks and precisely identifies a small number of genuinely memorized samples, thereby substantially enhancing the statistical rigor and reliability of memorization auditing for generative models.
📝 Abstract
Memorization audits of generative models read similarity scores against thresholds, with no null distribution, and the conclusions they support can be wrong. By MemBench's rule, the benchmark's mitigations roughly halve Stable Diffusion's memorization; audited with false-discovery control, two thirds of the certified images are no longer detected under random prompt perturbations, five sixths under attention rescaling, and all of them under embedding optimization. The field's data-copying test, read against its own null, flags ten of twenty-four generators that reproduce nothing. We argue that for memorization the null is the hard part, and supply two. For a whole model, training and held-out images are exchangeable given its samples, and relabelling them is a permutation test, exact for any statistic when the held-out images are a random split; under it, a nearest-neighbour preference still fires on seven of those twenty-four, and a count restricted to the near-duplicate scale on none (McNemar p=0.016). For single images, the natural nulls fail twice, measurably: ranking an image among random images yields 596 false discoveries among 2,365 controls, and resampling independent generations makes the null three times too narrow. Calibrated against matched controls, the audit certifies 36 of 61 MemBench images at 5% false-discovery rate, held-out controls are certified in 0.01% of calibration splits, and on this benchmark two generations per image recover that count. A calibrated maximum, which reads occasional rather than typical copying, certifies 46. As the scale-restricted statistic we recommend the small-scale mass of the Intersection Euler Characteristic Profile, which also counts distinct images copied and tests whether two models copy the same ones.