Compositional Concept Erasure in Text-to-Image Diffusion Models via Hierarchically Grounded Semantic Surgery

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the failure of hierarchical routing and attribute leakage in text-to-image diffusion models during concept erasure under compositional prompts. We propose a Hierarchical Semantic Surgery framework that introduces a novel span-level localization mechanism integrating lexical, categorical, and semantic evidence. Combined with a dynamic attribute binding algorithm and a cross-attention objective optimized via counterfactual references, our method achieves precise concept removal without weight updates. Experiments demonstrate that the proposed approach reduces the hierarchical bypass rate to 10.02 and halves attribute leakage, achieving state-of-the-art neighbor exclusion and attribute preservation on the SEE and UnlearnCanvas benchmarks.
📝 Abstract
Removing copyrighted, unsafe, or user-specified concepts from a deployed text-to-image diffusion model is now a practical requirement. Weight-editing methods can suppress fixed targets, but they require per-target retraining and modify the model checkpoint. Training-free methods, on the other hand, are deployment-friendly, but they suffer from text-side routing failures on compositional prompts. In such prompts, the erase target may be invoked through a related class rather than its lexical name, and its modifiers may migrate onto preserved objects. This paper proposes Hierarchically Grounded Semantic Surgery (HGSS), a training-free framework for compositional concept erasure. The framework lifts both the routing signal and the edit operator used by text-side erasure. First, hierarchical span grounding resolves erase-target spans through lexical, taxonomic, and semantic evidence, while guarding against broad-hypernym and compound-head false positives. Second, dynamic attribute binding refines the text conditioning during early denoising via a counterfactual reference and a preserve-aware cross-attention objective, keeping surviving attribute-noun bindings intact. HGSS selectively removes the erase target without updating model weights or adding learned parameters. On SEE, HGSS cuts hierarchical evasion from 29.54 to 10.02 and roughly halves pairwise attribute leakage, achieving the best Neighbor E and AttrP scores among the reported erasure methods. On UnlearnCanvas, HGSS slightly improves the six-metric average over the matched Semantic Surgery baseline, reaching state-of-the-art.
Problem

Research questions and friction points this paper is trying to address.

concept erasure
text-to-image diffusion models
compositional prompts
training-free
attribute leakage
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compositional Concept Erasure
Training-free
Hierarchical Span Grounding
Dynamic Attribute Binding
Diffusion Models
💼 Related Jobs
No related jobs found.
C
Chen Dai
Department of Computer Science, Virginia Tech, Alexandria, Virginia, USA
G
Ganyu Zou
Department of Computer Science, Virginia Tech, Alexandria, Virginia, USA
Nathan Self
Nathan Self
Research Associate, Discovery Analytics Center, Virginia Tech
interactive visualizationweb development
K
Kevin Piper
Verisign, Inc., Reston, Virginia, USA
R
Ramachandra Rao Seethiraju
Verisign, Inc., Reston, Virginia, USA
K
Karthik Shyamsunder
Verisign, Inc., Reston, Virginia, USA
C
Chang-Tien Lu
Department of Computer Science, Virginia Tech, Alexandria, Virginia, USA
Naren Ramakrishnan
Naren Ramakrishnan
Thomas L. Phillips Professor, Virginia Tech
ForecastingMachine LearningComputational epidemiologyRecommender systemsVisual analytics