🤖 AI Summary
This study addresses the inherent difficulty in balancing target concept erasure with the preservation of semantically related concepts during concept forgetting in diffusion models. Through geometric analysis of activation spaces and rigorous theoretical proofs, this work formally defines concept regions, derives a theoretical lower bound on model degradation induced by representation overlap, and proposes a Pareto front-based evaluation criterion. The research reveals fundamental trade-off limits in concept forgetting arising from geometric overlaps in concept representations, theoretically proving that perfect forgetting is unattainable. Extensive multi-method benchmarking further confirms that existing approaches are universally constrained by such overlaps, failing to simultaneously achieve strong erasure and high-fidelity generation of neighboring concepts. These findings establish a principled theoretical foundation for understanding and advancing concept forgetting in generative models.
📝 Abstract
Concept unlearning in text-to-image diffusion models aims to suppress a target concept (e.g., \texttt{horse}) while preserving related but distinct content (e.g., \texttt{donkey}), yet existing methods either leak under indirect prompts or visibly degrade other concepts. We show that these failure modes stem from the geometry of concept representations rather than from any particular algorithm. Formalizing concepts as activation-space regions, we prove that the overlap between a target and other concepts lower-bounds the damage any robust erasure must inflict on them, with the trade-off scaling linearly in the degree of overlap. Across thirteen unlearning methods, including methods designed to preserve non-target concepts, no method achieves both strong erasure and strong neighbor preservation: STEREO nearly eliminates indirect leakage but cuts neighbor generation by more than 75\%, while sparse inference-time methods preserve neighbors but leak. Damage increases with our overlap measure, monotonically so for STEREO; the $\kappa$-scaling reproduces on SDXL, and neighbor-selective damage recurs on FLUX. Perfect unlearning is the wrong target for entangled concepts; methods should be evaluated on the Pareto frontier our theorem establishes.