🤖 AI Summary
This study addresses the limitation of existing machine unlearning methods for text-to-image models, where sensitivity to initial noise leads to probabilistic forgetting and insufficient adversarial robustness. It reveals that standard Gaussian sampling is fundamentally misaligned with unlearning objectives and proposes a novel adaptive concept-conditioned noise sampling framework. This approach dynamically focuses on regions yielding effective gradient updates and incorporates gradient weighting techniques to reinforce the forgetting process. Under both black-box and white-box adversarial evaluations, the proposed method reduces the regeneration rate of nudity concepts by an average of 67.2% and significantly suppresses attack success rates, while effectively preserving overall image generation quality and performance on non-target concepts.
📝 Abstract
Machine unlearning has emerged as a critical post-hoc safety measure to erase sensitive concepts from Text-to-Image (T2I) models without prohibitive retraining. However, we reveal that current state-of-the-art (SOTA) approaches are brittle due to a severe lack of robustness to noise initialization. We call this phenomenon ``probabilistic forgetting'': suppressed concepts re-emerge under specific random initial noise conditions, despite appearing unlearned on other initializations. We trace this failure to the misalignment between standard Gaussian sampling during unlearning and the unlearning objective. Since the target concept manifests only in specific initial noise regions throughout the unlearning phase, uniform random sampling yields sparse, uninformative gradient updates that fail to drive robust erasure. To overcome this issue, we propose an adaptive, concept-conditioned sampling strategy that dynamically concentrates gradient updates on regions where the target concept manifests, down-weighting uninformative areas. We integrate our framework with six distinct SOTA unlearning methods across four diffusion backbones and evaluate it across safety, object, and artistic-style unlearning, as well as under black-box and white-box adversarial attacks. Our method reduces the conditional nudity re-emergence rate across random initializations by 67.2% on average over four baselines and lowers attack success rates across both adversarial evaluations. Across concept domains, Adaptive Noise Sampling strengthens adversarial robustness and non-target retention while preserving competitive generative quality and target-erasure performance.