🤖 AI Summary
This study addresses the vulnerability of existing concept erasure methods to adversarial attacks, which arises from their exclusive focus on redirecting text mappings while leaving residual visual generation trajectories intact. To overcome this limitation, this work proposes a novel paradigm that shifts from text-mapping redirection to visual trajectory redirection. Specifically, it introduces a dual-branch mechanism that steers visual trajectories toward concept-removed endpoints, integrated with structure-preserving editing and joint optimization strategies for model fine-tuning. This approach effectively bridges the gap in thoroughly eliminating visual knowledge. Experimental results demonstrate that the proposed method significantly enhances robustness across style, celebrity, and nudity erasure tasks, reducing the maximum attack success rate to 0%–8% while effectively preserving general generation quality, thereby achieving robust concept erasure.
📝 Abstract
Concept erasure is essential for the safe deployment of text-to-image diffusion models, as they may reproduce harmful, copyrighted, or privacy-sensitive content learned from unconstrained large-scale data. Existing methods typically erase unwanted concepts while preserving general generation capability by redirecting target-related text-to-image mappings. However, recent studies show that erased models may still retain visual generative trajectories of target concepts, leaving them vulnerable to adversarial recovery attacks and revealing a fundamental gap between redirecting text-to-image mappings and truly removing visual knowledge. To bridge this gap, we propose VisualErase, a new paradigm that redirects concept-bearing visual generative trajectories toward explicitly defined concept-removed outcomes. To enable this redirection, we use structure-preserving image editing to construct content-aligned, concept-removed counterparts for source images, providing explicit visual endpoints that retain non-target content. We then derive a denoising target from each source-to-counterpart pair and use a dual-branch redirection loss to align both text-conditioned and unconditional predictions with this target, since conditional supervision alone does not explicitly constrain generation without textual guidance. To mitigate the adverse effects of concept erasure on non-target generation, we jointly optimize the redirection loss with a counterpart retention loss that matches denoising predictions from the frozen pretrained model. Across style, celebrity, and nudity erasure, VisualErase limits the maximum attack success rate over seven attacks to 0%, 8%, and 0.1%, respectively, while retaining general generation quality. These results highlight the importance of visual trajectory redirection for robust concept erasure beyond text-to-image mappings alone.