🤖 AI Summary
Text-to-image (T2I) diffusion models frequently generate sensitive, harmful, or copyright-protected content, necessitating controllable concept suppression mechanisms. To address this, we systematically survey existing concept removal methods and introduce the first three-dimensional taxonomy—spanning intervention level, optimization architecture, and semantic scope—that exposes fundamental trade-offs among specificity, generalization, and efficiency. We propose an evaluation gap analysis framework and establish the first unified taxonomy and benchmark specifically designed for ethical alignment. By integrating gradient-driven editing, latent-space intervention, and concept disentanglement techniques, we empirically validate fine-grained semantic masking while preserving generation fidelity. Our contributions provide a reproducible, rigorously evaluable methodological foundation for responsible generative AI.
📝 Abstract
Text-to-Image (T2I) models have demonstrated impressive capabilities in generating high-quality and diverse visual content from natural language prompts. However, uncontrolled reproduction of sensitive, copyrighted, or harmful imagery poses serious ethical, legal, and safety challenges. To address these concerns, the concept erasure paradigm has emerged as a promising direction, enabling the selective removal of specific semantic concepts from generative models while preserving their overall utility. This survey provides a comprehensive overview and in-depth synthesis of concept erasure techniques in T2I diffusion models. We systematically categorize existing approaches along three key dimensions: intervention level, which identifies specific model components targeted for concept removal; optimization structure, referring to the algorithmic strategies employed to achieve suppression; and semantic scope, concerning the complexity and nature of the concepts addressed. This multi-dimensional taxonomy enables clear, structured comparisons across diverse methodologies, highlighting fundamental trade-offs between erasure specificity, generalization, and computational complexity. We further discuss current evaluation benchmarks, standardized metrics, and practical datasets, emphasizing gaps that limit comprehensive assessment, particularly regarding robustness and practical effectiveness. Finally, we outline major challenges and promising future directions, including disentanglement of concept representations, adaptive and incremental erasure strategies, adversarial robustness, and new generative architectures. This survey aims to guide researchers toward safer, more ethically aligned generative models, providing foundational knowledge and actionable recommendations to advance responsible development in generative AI.