Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair

šŸ“… 2026-10-01
šŸ“ˆ Citations: 0
✨ Influential: 0
šŸ“„ PDF
šŸ¤– AI Summary
This study addresses the unclear mechanisms underlying self-repair in language models by proposing an "ablation-as-dosage" perspective and a counterfactual contrast intensity coordinate system. Through fine-grained unit intervention experiments and affine regression analyses on cross-family large models including Gemma and LLaMA, we reveal that self-repair fundamentally constitutes pre-existing gain rather than dynamic adjustment. Component responses follow an affine law whose slope is predictable from fixed weights, and seemingly anomalous phenomena reflect the standard operation of "anti-gravity units" under contrastive signals. Experiments demonstrate that 84% of downstream directions across four models conform to this law, and seven anti-gravity attention heads are precisely identified within the GPT-2 IOI circuit. This work provides the first unified explanation for self-repair noise and establishes linear regularities governing causal repair.
šŸ“ Abstract
Ablate a component of a language model, and other components often appear to adjust and compensate. This phenomenon, termed self-repair, has been observed repeatedly, but its mechanism remains unclear. The most systematic study to date concluded that self-repair is noisy and unlikely to have a single explanation. We argue that it has one: a gain already present before any ablation. Any intervention on a causally important component can be viewed as a point on a coordinate axis $λ$, the signed strength of a counterfactual contrast. Hence, conventional ablation methods are uncalibrated points on this axis. We show that the causal repair response for a fine-grained unit $r$ is governed by an affine law, $E_r(λ)=\mathrm{own}_r+γ_rλ$. The slope $γ_r$ is a fixed coefficient that consistently influences the model, with or without ablation, and its sign determines whether the unit counteracts or reinforces the removed signal. On a factual-verdict task across four models from distinct families (Gemma, Qwen, LLaMA, and Mistral), we identify components including MLP neurons, OV neurons, and singular directions that follow this affine law, 68 of 81 downstream directions in all. Moreover, we can anticipate the magnitude of $γ_r$ from the fixed weights. On the IOI circuit of GPT-2 Small, seven of the ten heads the intervention can reach follow the law, and all seven are counterweights. From this perspective, what may appear as self-repair is a counterweight performing its usual operation when the contrastive signal emerges at the core.
Problem

Research questions and friction points this paper is trying to address.

self-repair
ablation
language models
mechanistic interpretability
counterweights
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-repair
Ablation
Affine law
Counterweights
Causal intervention
šŸ”Ž Similar Papers
No similar papers found.