LEACE: Perfect linear concept erasure in closed form

📅 2023-06-06
🏛️ Neural Information Processing Systems
📈 Citations: 96
✨ Influential: 14
📄 PDF
🤖 AI Summary
This work addresses the problem of controllably removing specific semantic attributes—such as gender, race, or part-of-speech—from neural embeddings. We propose LEACE (Linear Exact Adversarial Concept Erasure), the first method to yield a provably optimal closed-form linear erasure solution: it strictly guarantees that *no* linear classifier can detect the target concept while minimizing ℓ₂ embedding perturbation. LEACE achieves this via orthogonal projection under covariance constraints and joint optimization across layers, enabling end-to-end “concept scrubbing” in large language models. Experiments on BERT demonstrate that LEACE significantly reduces gender bias, drops part-of-speech prediction accuracy to chance level (≈50%), and incurs the smallest ℓ₂ distortion among baselines. Crucially, it is the first linear intervention achieving cross-layer, verifiably fair, and information-preserving concept control—i.e., fairness guarantees hold provably without sacrificing downstream utility.
📝 Abstract
Concept erasure aims to remove specified features from an embedding. It can improve fairness (e.g. preventing a classifier from using gender or race) and interpretability (e.g. removing a concept to observe changes in model behavior). We introduce LEAst-squares Concept Erasure (LEACE), a closed-form method which provably prevents all linear classifiers from detecting a concept while changing the embedding as little as possible, as measured by a broad class of norms. We apply LEACE to large language models with a novel procedure called"concept scrubbing,"which erases target concept information from every layer in the network. We demonstrate our method on two tasks: measuring the reliance of language models on part-of-speech information, and reducing gender bias in BERT embeddings. Code is available at https://github.com/EleutherAI/concept-erasure.
Problem

Research questions and friction points this paper is trying to address.

Erasing specified features from embeddings to improve fairness
Preventing linear classifiers from detecting target concepts
Reducing gender bias in BERT embeddings via concept scrubbing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Closed-form least-squares concept erasure
Concept scrubbing for all network layers
Minimizes embedding change under norms
🔎 Similar Papers
No similar papers found.