Neuralyzing the Trace: Selective Representation-Level Unlearning with Contrastive Sparse Autoencoders

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of precisely localizing and removing specific internal representations during privacy-oriented machine unlearning in large language models. To this end, it proposes SCALPEL (Sparse Contrastive Autoencoder for Precise Erasure Learning), a method that exploits reconstruction energy deviations and introduces a contrastive training strategy to enhance selective forgetting of target features. By achieving representation-level de-knowledging with theoretically guaranteed controllability over contextual perturbation, SCALPEL integrates mechanistic interpretability with representation intervention techniques. Empirical evaluations on the TOFU benchmark demonstrate that SCALPEL significantly outperforms NMF and standard SAEs while matching the performance of gradient difference and RMU methods. This work establishes a new paradigm for precise model unlearning, offering both theoretical guarantees and strong empirical effectiveness.
📝 Abstract
Machine unlearning aims to remove targeted information while preserving a model's other abilities. In realistic settings, such as privacy requests under the EU GDPR, the target may be narrow, for example information associated with a single person. Behavioral forgetting alone may be insufficient, motivating interventions directly on internal representations. However, standard mechanistic-interpretability extractors are poorly selective for such targets. We identify an energy bias in reconstruction-based extraction, which favors dominant background structure over low-energy target-specific components. We introduce SCALPEL, a contrastive sparse autoencoder designed to learn more selective forget features. We show theoretically that contrastive training promotes target-selective features and that our selection score controls expected background knowledge perturbation. We validate SCALPEL experimentally on TOFU across Qwen, Llama, and Gemma, where it substantially improves over NMF and standard SAE interventions and is competitive with Gradient Difference and RMU, bridging mechanistic interpretability and fine-grained unlearning.
Problem

Research questions and friction points this paper is trying to address.

Machine Unlearning
Representation-Level Forgetting
Mechanistic Interpretability
Sparse Autoencoders
Energy Bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

Machine Unlearning
Contrastive Sparse Autoencoder
Mechanistic Interpretability
Representation-Level Intervention
Energy Bias