You Can't Have It Both Ways: Concept Entanglement Limits Diffusion Model Unlearning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent difficulty in balancing target concept erasure with the preservation of semantically related concepts during concept forgetting in diffusion models. Through geometric analysis of activation spaces and rigorous theoretical proofs, this work formally defines concept regions, derives a theoretical lower bound on model degradation induced by representation overlap, and proposes a Pareto front-based evaluation criterion. The research reveals fundamental trade-off limits in concept forgetting arising from geometric overlaps in concept representations, theoretically proving that perfect forgetting is unattainable. Extensive multi-method benchmarking further confirms that existing approaches are universally constrained by such overlaps, failing to simultaneously achieve strong erasure and high-fidelity generation of neighboring concepts. These findings establish a principled theoretical foundation for understanding and advancing concept forgetting in generative models.
📝 Abstract
Concept unlearning in text-to-image diffusion models aims to suppress a target concept (e.g., \texttt{horse}) while preserving related but distinct content (e.g., \texttt{donkey}), yet existing methods either leak under indirect prompts or visibly degrade other concepts. We show that these failure modes stem from the geometry of concept representations rather than from any particular algorithm. Formalizing concepts as activation-space regions, we prove that the overlap between a target and other concepts lower-bounds the damage any robust erasure must inflict on them, with the trade-off scaling linearly in the degree of overlap. Across thirteen unlearning methods, including methods designed to preserve non-target concepts, no method achieves both strong erasure and strong neighbor preservation: STEREO nearly eliminates indirect leakage but cuts neighbor generation by more than 75\%, while sparse inference-time methods preserve neighbors but leak. Damage increases with our overlap measure, monotonically so for STEREO; the $\kappa$-scaling reproduces on SDXL, and neighbor-selective damage recurs on FLUX. Perfect unlearning is the wrong target for entangled concepts; methods should be evaluated on the Pareto frontier our theorem establishes.
Problem

Research questions and friction points this paper is trying to address.

concept unlearning
diffusion models
concept entanglement
activation-space overlap
Pareto frontier
Innovation

Methods, ideas, or system contributions that make the work stand out.

Concept Unlearning
Diffusion Models
Representation Geometry
Pareto Frontier
Concept Entanglement
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yian Wang
University of Illinois Urbana-Champaign
A
Ali Ebrahimpour-Boroojeny
University of Illinois Urbana-Champaign
H
Hari Sundaram
University of Illinois Urbana-Champaign
Varun Chandrasekaran
Varun Chandrasekaran
University of Illinois Urbana-Champaign
SecurityPrivacyArtificial Intelligence