edit token embeddings

Designs and evaluates methods that modify token embedding vectors to remove, alter, or erase concept-specific components at the embedding level, by changing only the affected token vectors while minimizing coherence degradation across other tokens; includes techniques for selective embedding erasure and analyses of how easily the removed information can be recovered via relearning.

edittokenembeddings

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Erased or Dormant? Rethinking Concept Erasure Through Reversibility

May 22, 2025
PL
Ping Liu
🏛️ University of Nevada Reno | National University of Singapore

This work investigates whether concept erasure in text-to-image diffusion models genuinely eliminates the model’s capacity to generate a target concept or merely achieves prompt-dependent, superficial suppression. To address this, we propose the first instance-level reversibility evaluation paradigm, which employs lightweight fine-tuning (<100 steps) to probe whether erased concepts can be reactivated across diverse prompts. Experimental results demonstrate that mainstream erasure methods preserve underlying semantic representations: most “erased” concepts are robustly and faithfully reinstated under cross-prompt conditions. This confirms that current techniques implement reversible suppression—effectively placing concepts into a dormant state—rather than irreversible removal. Our findings shift the focus of concept editing from parameter-space modification toward representation-level, irreversible interventions. The proposed evaluation framework establishes a new benchmark for assessing conceptual integrity in generative models and provides a principled technical pathway for trustworthy AI content governance.

Assess if concept erasure truly removes generative capacity in diffusion modelsEvaluate robustness and reversibility of current concept erasure techniquesTest reactivation potential of erased concepts through fine-tuning

This work addresses a critical oversight in existing knowledge erasure methods—the frequent neglect of the embedding layer—which renders erased knowledge vulnerable to recovery via adversarial prompts or relearning. The study is the first to highlight the pivotal role of the embedding layer in effective knowledge removal and introduces EMBER, a plug-and-play, embedding-level intervention module. EMBER leverages sparse matrix factorization to precisely identify and edit word embeddings associated with target concepts. It seamlessly integrates with existing parameter-update-based erasure techniques and significantly enhances both robustness and specificity of erasure on Gemma-2-2B-it and Llama-3.1-8B-Instruct models: relearning-based recovery accuracy is reduced by up to 50%, remaining below 35%, while preserving linguistic coherence for nearly all tokens except a minimal set of concept-specific words.

embedding layerknowledge erasurelanguage models

EraseBench: Understanding The Ripple Effects of Concept Erasure Techniques

Jan 16, 2025
IA
Ibtihel Amara
🏛️ Google Research | McGill University | Rice University | Google Deepmind

This work systematically exposes the cascading failure of “concept erasure” techniques in text-to-image models when handling visually similar, semantically associated, or binary-opposite concepts—termed “concept confusion” and the newly identified “concept ripple effect.” To enable rigorous quantification, we introduce EraseBENCH, the first benchmark tailored for multi-dimensional concept relationships, comprising 100+ concepts and 1,000+ curated prompts. We propose a unified evaluation framework integrating multi-faceted prompt engineering, concept relationship modeling, and adversarial interference testing, with dual metrics assessing both image quality and semantic fidelity. Extensive experiments reveal that state-of-the-art erasure methods suffer from substantial degradation in visual quality and pervasive semantic leakage under realistic conditions, undermining their reliability for industrial deployment.

Concept ConfusionConcept ErasureImage Generation Models

Robust Concept Erasure Using Task Vectors

Apr 04, 2024
MP
Minh Pham
🏛️ New York University

This work addresses the critical safety challenge of **unconditional and robust erasure of harmful concepts** in text-to-image models—without requiring user-provided prompts or interventions. Methodologically, it introduces (1) **Diverse Inversion**, a latent-space inversion strategy that enhances diversity and robustness in concept representation estimation; (2) a **concept-driven embedding set construction** coupled with **sparse weight subset editing**, enabling precise localization and attenuation of parameters associated with the target concept; and (3) **Task Vectors** to dynamically estimate optimal editing strength. Experiments demonstrate that the method significantly improves generalization and robustness of concept erasure under unseen prompts, drastically reduces unintended concept deletion, and preserves over 92% of the original model’s generation quality and diversity. The approach provides a scalable, model-level defense mechanism for safe and controllable deployment of generative AI systems.

Maintaining core model performance during erasureRobustness to unexpected user inputsUnconditional concept erasure in text-to-image models

LEACE: Perfect linear concept erasure in closed form

Jun 06, 2023
NB
Nora Belrose
🏛️ EleutherAI | Bar-Ilan University | ETH Zürich | Booz Allen Hamilton

This work addresses the problem of controllably removing specific semantic attributes—such as gender, race, or part-of-speech—from neural embeddings. We propose LEACE (Linear Exact Adversarial Concept Erasure), the first method to yield a provably optimal closed-form linear erasure solution: it strictly guarantees that *no* linear classifier can detect the target concept while minimizing ℓ₂ embedding perturbation. LEACE achieves this via orthogonal projection under covariance constraints and joint optimization across layers, enabling end-to-end “concept scrubbing” in large language models. Experiments on BERT demonstrate that LEACE significantly reduces gender bias, drops part-of-speech prediction accuracy to chance level (≈50%), and incurs the smallest ℓ₂ distortion among baselines. Crucially, it is the first linear intervention achieving cross-layer, verifiably fair, and information-preserving concept control—i.e., fairness guarantees hold provably without sacrificing downstream utility.

Erasing specified features from embeddings to improve fairnessPreventing linear classifiers from detecting target conceptsReducing gender bias in BERT embeddings via concept scrubbing

Latest Papers

What's happening recently
View more

This work addresses a critical gap in existing concept erasure methods for text-to-video (T2V) diffusion models, which only verify the absence of target concepts in output frames without assessing whether internal representations are truly removed. To this end, we propose PROBE, a diagnostic protocol that quantifies the reactivation potential of erased concepts by optimizing lightweight pseudo-token embeddings under latent alignment constraints while keeping model parameters frozen. We introduce a multi-level evaluation framework—integrating classifier-based detection, semantic similarity metrics, temporal reactivation analysis, and human validation—and reveal, for the first time, that current approaches achieve only output-level suppression. Our experiments across three T2V architectures, three concept categories, and three erasure strategies demonstrate measurable residual concept capacity in all cases, with robustness strongly correlated to the depth of temporal intervention.

concept erasurereactivation potentialresidual capacity

This work addresses the issue that existing training-free methods for erasing specific concepts from diffusion models often inadvertently remove semantically related non-target content. To mitigate this, the authors propose CARE, a closed-form concept erasure operator that constructs a perceptually preserved subspace in the cross-attention value space, guided by anchor representations of retained concepts. The target concept direction is then replaced with its projection onto this subspace, enabling precise erasure while preserving shared visual structures. CARE incorporates an adjustable shrinkage parameter to balance erasure efficacy and semantic retention, and it offers theoretical guarantees of minimal perturbation. Experiments demonstrate that CARE significantly outperforms current state-of-the-art methods across instance-, style-, and celebrity-level concept erasure tasks, while effectively safeguarding unrelated semantic information.

collateral damageconcept erasurecross-attention

This work addresses the challenge that existing concept erasure methods struggle to simultaneously achieve robust erasure and high-fidelity generation of non-target concepts. To this end, the authors propose PARSE, a training-free framework that performs preservation-aware concept erasure in the cross-attention value space. Its key innovations include classifier-free guidance–based dynamic token-level concept discovery, preservation-aware subspace projection, adaptive subspace expansion, and textual inversion trigger search. The paper also introduces BEUS, a comprehensive evaluation metric that balances attack success rate against generation quality. Experiments demonstrate that PARSE significantly outperforms current approaches on NSFW, artistic style, and object erasure tasks, achieving robust multi-concept removal while preserving high-fidelity image generation.

concept erasurediffusion modelsmodel editing

Existing text-guided diffusion models for concept erasure are constrained by the representational capacity of textual space and rely on predefined erasure references, often compromising overall generation performance. This work proposes a component-level disentangled forgetting mechanism that decomposes concept embeddings into critical and non-critical components via a Component Extraction Module (CEM) and a Swap Disentanglement Strategy (SDS). By selectively removing only the harmful portions and fine-tuning weights accordingly—without requiring any predefined reference—the method achieves precise and robust concept erasure. It effectively forgets target content while substantially preserving the model’s general generative capabilities, outperforming current alignment-based fine-tuning approaches.

concept unlearningdiffusion modelsharmful content generation

This work proposes High-order Semantic Representation Misdirection (HiRM), a method to prevent the misuse of text-to-image diffusion models for generating harmful, privacy-sensitive, or copyrighted content by precisely erasing specific concepts without degrading the generation quality of unrelated ones. By leveraging causal tracing, HiRM identifies visual attribute representations of target concepts in early self-attention layers of the text encoder and redirects them toward random or semantic superclass directions. Only these critical layers are fine-tuned, yielding an efficient, decoupled erasure mechanism that operates independently of the denoiser. Evaluated on UnlearnCanvas and NSFW benchmarks, HiRM effectively removes diverse targets—including objects, artistic styles, and nudity—while preserving high image fidelity and incurring low training costs. Notably, the approach demonstrates zero-shot transferability to advanced architectures such as Flux.

concept erasurelocalized representationmisuse mitigation

Hot Scholars

PH

Ping He

Zhejiang University
AI Security
TS

Teng Shi

Renmin University of China
Recommender SystemInformation Retrieval
JX

Jun Xu

Professor, Gaoling School of Artificial Intelligence, Renmin University of China
Information RetrievalLearning to RankSemantic Matching