Erased or Dormant? Rethinking Concept Erasure Through Reversibility

📅 2025-05-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates whether concept erasure in text-to-image diffusion models genuinely eliminates the model’s capacity to generate a target concept or merely achieves prompt-dependent, superficial suppression. To address this, we propose the first instance-level reversibility evaluation paradigm, which employs lightweight fine-tuning (<100 steps) to probe whether erased concepts can be reactivated across diverse prompts. Experimental results demonstrate that mainstream erasure methods preserve underlying semantic representations: most “erased” concepts are robustly and faithfully reinstated under cross-prompt conditions. This confirms that current techniques implement reversible suppression—effectively placing concepts into a dormant state—rather than irreversible removal. Our findings shift the focus of concept editing from parameter-space modification toward representation-level, irreversible interventions. The proposed evaluation framework establishes a new benchmark for assessing conceptual integrity in generative models and provides a principled technical pathway for trustworthy AI content governance.

Technology Category

Computer Vision: Diffusion Models for VisionCognitive Modeling & Cognitive Systems: Conceptual Inference and ReasoningNatural Language Processing: Generation

Application Category

Search and Retrieval-Augmented AI: Web query analysis, representation and understandingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
To what extent does concept erasure eliminate generative capacity in diffusion models? While prior evaluations have primarily focused on measuring concept suppression under specific textual prompts, we explore a complementary and fundamental question: do current concept erasure techniques genuinely remove the ability to generate targeted concepts, or do they merely achieve superficial, prompt-specific suppression? We systematically evaluate the robustness and reversibility of two representative concept erasure methods, Unified Concept Editing and Erased Stable Diffusion, by probing their ability to eliminate targeted generative behaviors in text-to-image models. These methods attempt to suppress undesired semantic concepts by modifying internal model parameters, either through targeted attention edits or model-level fine-tuning strategies. To rigorously assess whether these techniques truly erase generative capacity, we propose an instance-level evaluation strategy that employs lightweight fine-tuning to explicitly test the reactivation potential of erased concepts. Through quantitative metrics and qualitative analyses, we show that erased concepts often reemerge with substantial visual fidelity after minimal adaptation, indicating that current methods suppress latent generative representations without fully eliminating them. Our findings reveal critical limitations in existing concept erasure approaches and highlight the need for deeper, representation-level interventions and more rigorous evaluation standards to ensure genuine, irreversible removal of concepts from generative models.
Problem

Research questions and friction points this paper is trying to address.

Assess if concept erasure truly removes generative capacity in diffusion models
Evaluate robustness and reversibility of current concept erasure techniques
Test reactivation potential of erased concepts through fine-tuning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evaluates robustness and reversibility of concept erasure
Uses lightweight fine-tuning to test reactivation potential
Proposes representation-level interventions for irreversible removal
🔎 Similar Papers