🤖 AI Summary
This study addresses the limitation of existing concept erasure evaluations, which rely on architecture-specific attacks and lack a unified robustness testing methodology across paradigms. We propose the CSR framework, which achieves cross-architecture unified concept reactivation without requiring external image data by optimizing directly within the native parameter spaces of U-Net and Transformer architectures. By integrating parameter-level optimization with prediction space mapping, CSR seamlessly adapts to both Stable Diffusion and FLUX models. Experimental results demonstrate that under strict nudity evaluation settings, CSR attains average attack success rates of 50.47% and 40.29% on FLUX and SD, respectively, achieving state-of-the-art performance. These findings reveal the persistent recoverability of supposedly erased concepts, underscoring critical vulnerabilities in current concept erasure techniques.
📝 Abstract
Concept erasure aims to suppress undesirable knowledge in text-to-image generative models. However, existing robustness evaluations typically rely on relearning attacks tailored to specific model architectures. We study concept reactivation across two substantially different generative paradigms: noise-prediction U-Nets and flow-matching Transformers. We introduce \textbf{Concept Score Relearning (CSR)}, a unified parameter-level framework that reactivates erased concepts by optimizing each model within its native prediction space. CSR requires no external target-concept image dataset and applies the same concept-directed objective to both U-Net-based Stable Diffusion and Transformer-based FLUX. Experiments across diverse concepts and multiple erasure methods demonstrate consistent concept reactivation across both architectures, highlighting the cross-architecture applicability of CSR and the persistent recoverability of apparently erased concepts. For strict nudity, CSR reaches average ASRs of 50.47\% on FLUX and 40.29\% on Stable Diffusion, consistently ranking first across all evaluated safety settings.