Causal Routing for Unlearning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive computational overhead of full-weight rewriting and the difficulty of localizing concept distributions in large language model unlearning by proposing a precise forgetting method based on causal routing. Specifically, target neurons are identified by evaluating activation variability through a single forward pass, and a parameter-efficient routing module is introduced to dynamically gate and suppress relevant neurons during inference. This work pioneers a query-time intervention paradigm that keeps the base model frozen, ensuring that behavioral modifications are exclusively attributable to controlled neurons while enhancing interpretability. Experimental results demonstrate that memory consumption is reduced to 14 GiB, performance retention on TOFU exhibits no significant deviation from the original model, and the adversarial probe recall on RWKU decreases to 0.052.
📝 Abstract
LLMs cannot forget the way we delete a file. Strangely, we are asked to remove something that was never put anywhere in particular. What the model took from a piece of text is now smeared across billions of weights. Existing methods rewrite all of them to change one thing, and none of them say which part produced that change. To address this, we introduce Causal Routing for Unlearning (CRU) by asking where the concept is expressed in the model and suppressing only that part. One untrained forward pass over the forget set ranks neurons by how their activations vary. Then, small routing modules on those neurons gate and suppress only the concepts that need to be forgotten. In CRU, the base model is frozen, and any change in behavior is caused only by the gated neurons; hence, why the routing is causal. Due to our parameter efficiency (only ~0.01% as many parameters as the base model), unlearning a concept costs 14 GiB, whereas the baselines require 71 GiB. On TOFU, CRU is indistinguishable from the retained model (p > 0.05, KS test) and is never Pareto-dominated, whereas every compared baseline matches its forgetting on the larger-forget batches only by collapsing utility. On RWKU, it achieves an adversarial-probe recall of 0.052, compared to 0.250 for the strongest baseline, meaning the knowledge is gone, not merely harder to reach. Thus, deciding on the intervention at query time, rather than fixing it beforehand, is the axis along which we argue that unlearning should proceed.
Problem

Research questions and friction points this paper is trying to address.

Machine Unlearning
Large Language Models
Knowledge Forgetting
Utility Preservation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Causal Routing
Machine Unlearning
Neuron Gating
Parameter Efficiency
Query-time Intervention
🔎 Similar Papers
No similar papers found.