No Concept Escapes the Audit: Auditing-Aware Unlearning for Verifiable Concept Erasure in Diffusion Models

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the security vulnerability in diffusion models whereby prohibited content remains recoverable from latent representations even after concept erasure. To mitigate this, we propose AVCE, a verifiable concept erasure framework that, for the first time, anchors erasure at geometrically weakest points. The method achieves thorough removal by auditing embedding neighborhoods and applying closed-form editing to cross- and self-attention projections. Furthermore, it combines orthogonal gradient projection with path-level audit loss fine-tuning, thereby overcoming the limitations of conventional text-interface interventions. Experimental evaluations on models such as Stable Diffusion v1.5 demonstrate that AVCE reduces attack success rates by 5.07× and improves audit scores by 3.84× while preserving generation quality without degradation.
📝 Abstract
Text-to-image diffusion models can generate prohibited content, which motivates concept erasure through machine unlearning. Most erasure methods intervene at the text interface, through prompt modification or localized updates to text-conditioning weights, and they are evaluated by what the model outputs for given prompts. Such evaluation cannot see what the network still encodes. Latent-space auditing, which bypasses text conditioning and probes the denoising network directly, shows that erased concepts remain recoverable from internal representations. We find that this also holds for methods built to be robust against adversarial prompts, and that the problem grows with the number of erased concepts. We propose Auditing-Aware Unlearning for Verifiable Concept Erasure in Diffusion Models (AVCE), a framework that grounds erasure in the model's latent representations. AVCE audits the embedding neighborhood of each concept and condenses the discovered vulnerable directions into an anchor at the weakest geometric point. It edits cross-attention and self-attention projections in closed form at this anchor, then fine-tunes the two pathways with pathway-level auditing losses, using orthogonal gradient projection to consolidate multiple concepts. Experiments on SD v1.5, SDXL, and Flux 1.0 across object, explicit-content, and artistic-style unlearning show that AVCE reduces attack success rates by 5.07x and improves auditing scores by 3.84x over the strongest baseline, while preserving competitive generation quality.
Problem

Research questions and friction points this paper is trying to address.

concept erasure
machine unlearning
diffusion models
latent-space auditing
verifiable erasure
Innovation

Methods, ideas, or system contributions that make the work stand out.

Machine Unlearning
Latent-space Auditing
Diffusion Models
Concept Erasure
Orthogonal Gradient Projection
🔎 Similar Papers