🤖 AI Summary
State-of-the-art image generation models, trained on large-scale unfiltered datasets, are prone to producing undesirable visual concepts, and effectively removing such content without degrading generation quality remains highly challenging. This work proposes a novel integrated filtering mechanism that replaces the model’s bottleneck layer with a transcoder capable of structuring feature activations, enabling selective suppression of specific conceptual signals directly within the backbone network. For the first time, this approach achieves persistent, white-box, and robust concept removal inside the model backbone, supporting sequential deletion of multiple concepts without requiring external components. It is applicable across diverse architectures—including diffusion models like SD3.5 and Flux, as well as autoregressive models such as Infinity. Experiments demonstrate that the method attains state-of-the-art removal efficacy while preserving high-fidelity generation, exhibiting strong robustness against adversarial prompts and effective multi-concept handling capabilities.
📝 Abstract
Image generative models are trained on massive, largely uncurated internet-scale datasets that contain undesirable visual concepts. Efficiently removing such concepts from the model generations without degrading the quality of output images remains challenging. We introduce a novel concept removal method for frontier diffusion and image autoregressive models, such as SD3.5, Flux, and Infinity. Our intervention replaces the internal bottleneck layer present in all these modern models with a transcoder that is trained to replicate the original layer while structuring it into distinct activation features. This in-place substitution creates an integrated filter through which concept-specific signals can be selectively disabled while preserving the rest of the model's behavior. Since the intervention modifies the model backbone rather than attaching an external component, it remains persistent under white-box access. Empirically, the approach achieves state-of-the-art concept removal performance across modern diffusion and autoregressive models, maintains visual generation quality, provides robustness against adversarial prompts, and supports sequential removal of diverse concepts. This positions our method as a practical approach for concept removal in frontier image generative models.