🤖 AI Summary
This study addresses the unintended degradation of preserved concepts in diffusion model concept erasure, caused by representational overlap between target and retained concepts. To this end, we propose RASteer, a training-free, retention-aware activation steering method. Specifically, RASteer introduces a novel Retention Orthogonal Steering (ROS) mechanism that constructs a preservation subspace and orthogonalizes the erasure direction to enable precise concept removal. Furthermore, an Overlap-Adaptive Calibration (OAC) module is incorporated to dynamically balance erasure intensity against content fidelity during inference. Extensive experiments demonstrate that RASteer significantly outperforms existing baselines across multiple benchmarks, effectively enhancing both concept erasure precision and the generation quality of non-target content.
📝 Abstract
Concept erasure aims to remove a target concept, such as a copyrighted style, a recognizable character, or unsafe content, from a pretrained text-to-image diffusion model while preserving its ability to generate other content. Existing activation steering methods build an erasure direction mainly from the target concept and adjust model activations along it at inference time. However, target and retained concepts often overlap in the model's representation space, so this direction also contains shared components that retained concepts rely on. Steering directly along this direction can therefore suppress retained concepts and harm the generation of non-target content. To address this issue, we propose Retain-aware Activation Steering (RASteer), a training-free method. RASteer first builds a retain subspace from the concepts to preserve. Retain-Orthogonal Steering (ROS) then removes components aligned with this subspace from the erasure direction, making steering more specific to the target. Since fully removing the shared components can weaken erasure, we further introduce Overlap-Adaptive Calibration (OAC). At each layer and denoising step, OAC uses the overlap between the erasure direction and the retain subspace to control how much of each shared component is removed, balancing target erasure and concept preservation. Experiments on unsafe-content, instance, and artistic-style erasure across multiple backbones and benchmarks show that RASteer matches or outperforms the activation steering and weight editing baselines we evaluate, achieving a better balance between erasure and preservation.