Weeding Out Bad Seeds: Initial-Noise-Robust Unlearning for Text-to-Image Diffusion Models

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing machine unlearning methods for text-to-image models, where sensitivity to initial noise leads to probabilistic forgetting and insufficient adversarial robustness. It reveals that standard Gaussian sampling is fundamentally misaligned with unlearning objectives and proposes a novel adaptive concept-conditioned noise sampling framework. This approach dynamically focuses on regions yielding effective gradient updates and incorporates gradient weighting techniques to reinforce the forgetting process. Under both black-box and white-box adversarial evaluations, the proposed method reduces the regeneration rate of nudity concepts by an average of 67.2% and significantly suppresses attack success rates, while effectively preserving overall image generation quality and performance on non-target concepts.
📝 Abstract
Machine unlearning has emerged as a critical post-hoc safety measure to erase sensitive concepts from Text-to-Image (T2I) models without prohibitive retraining. However, we reveal that current state-of-the-art (SOTA) approaches are brittle due to a severe lack of robustness to noise initialization. We call this phenomenon ``probabilistic forgetting'': suppressed concepts re-emerge under specific random initial noise conditions, despite appearing unlearned on other initializations. We trace this failure to the misalignment between standard Gaussian sampling during unlearning and the unlearning objective. Since the target concept manifests only in specific initial noise regions throughout the unlearning phase, uniform random sampling yields sparse, uninformative gradient updates that fail to drive robust erasure. To overcome this issue, we propose an adaptive, concept-conditioned sampling strategy that dynamically concentrates gradient updates on regions where the target concept manifests, down-weighting uninformative areas. We integrate our framework with six distinct SOTA unlearning methods across four diffusion backbones and evaluate it across safety, object, and artistic-style unlearning, as well as under black-box and white-box adversarial attacks. Our method reduces the conditional nudity re-emergence rate across random initializations by 67.2% on average over four baselines and lowers attack success rates across both adversarial evaluations. Across concept domains, Adaptive Noise Sampling strengthens adversarial robustness and non-target retention while preserving competitive generative quality and target-erasure performance.
Problem

Research questions and friction points this paper is trying to address.

Machine Unlearning
Text-to-Image Diffusion Models
Probabilistic Forgetting
Noise Initialization Robustness
Concept Erasure
Innovation

Methods, ideas, or system contributions that make the work stand out.

Machine Unlearning
Text-to-Image Diffusion Models
Probabilistic Forgetting
Adaptive Noise Sampling
Adversarial Robustness
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Arian Komaei Koma
Sharif University of Technology
S
Seyed Amir Kasaei
Sharif University of Technology
A
Aida Aryafar
Sharif University of Technology
M
Matin Ghiasi
Sharif University of Technology
A
Ali Aghayari
Hong Kong University of Science and Technology
A
Amirhossein Souri
Sharif University of Technology
M
Mohammad Mosayyebi
Sharif University of Technology
A
AmirMahdi Sadeghzadeh
Sharif University of Technology
Mohammad Hossein Rohban
Mohammad Hossein Rohban
Associate Professor in Computer Engineering, Sharif University of Technology
Machine LearningStatisticsComputational Biology