🤖 AI Summary
This work addresses the problem of controllably removing specific semantic attributes—such as gender, race, or part-of-speech—from neural embeddings. We propose LEACE (Linear Exact Adversarial Concept Erasure), the first method to yield a provably optimal closed-form linear erasure solution: it strictly guarantees that *no* linear classifier can detect the target concept while minimizing ℓ₂ embedding perturbation. LEACE achieves this via orthogonal projection under covariance constraints and joint optimization across layers, enabling end-to-end “concept scrubbing” in large language models. Experiments on BERT demonstrate that LEACE significantly reduces gender bias, drops part-of-speech prediction accuracy to chance level (≈50%), and incurs the smallest ℓ₂ distortion among baselines. Crucially, it is the first linear intervention achieving cross-layer, verifiably fair, and information-preserving concept control—i.e., fairness guarantees hold provably without sacrificing downstream utility.
📝 Abstract
Concept erasure aims to remove specified features from an embedding. It can improve fairness (e.g. preventing a classifier from using gender or race) and interpretability (e.g. removing a concept to observe changes in model behavior). We introduce LEAst-squares Concept Erasure (LEACE), a closed-form method which provably prevents all linear classifiers from detecting a concept while changing the embedding as little as possible, as measured by a broad class of norms. We apply LEACE to large language models with a novel procedure called"concept scrubbing,"which erases target concept information from every layer in the network. We demonstrate our method on two tasks: measuring the reliance of language models on part-of-speech information, and reducing gender bias in BERT embeddings. Code is available at https://github.com/EleutherAI/concept-erasure.