activation patching

Design and implement interventions that replace, edit, or transplant internal neural activations (e.g., overwriting with donor activations, mean vectors, or edited vectors, and performing back‑patching or single‑step interchange) and the tooling to apply those patches. Analyze how these activation-level interventions change model outputs to attribute behavior to components, estimate causal effects (including indirect effects), quantify intervention effectiveness across layers or authority levels, and test reversibility of erased representations.

activationpatching

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.18
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$214K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing activation intervention methods rely heavily on empirical design and lack theoretical grounding, making it difficult to systematically achieve efficient model adaptation. This work establishes, for the first time, a first-order equivalence between activation-space interventions and weight-space fine-tuning, revealing that the outputs of later transformer blocks serve as highly expressive intervention locations. Building on this insight, we propose a theoretically grounded strategy for selecting optimal intervention points. Furthermore, we introduce a novel weight-activation joint adaptation paradigm that simultaneously optimizes in both spaces. By training only 0.04% of the model parameters, our method achieves 99.1%–99.8% of the performance of full fine-tuning across multiple tasks, significantly outperforming mainstream parameter-efficient approaches such as ReFT and LoRA.

activation steeringintervention locationmodel adaptation

Composable Interventions for Language Models

Jul 09, 2024
AK
Arinbjörn Kolbeinsson
🏛️ University of Virginia | EleutherAI | Microsoft | University of Exeter | Eindhoven University of Technology | Harvard Medical School | University of Oxford | Mass General Brigham | UNC Chapel Hill

Existing test-time interventions—such as knowledge editing, model compression, and machine unlearning—have evolved in isolation, with no standardized framework to systematically study their interactions when applied jointly on the same model. Method: We propose the first composable intervention framework, unifying the modeling of synergistic mechanisms across these three intervention types. We design a comprehensive evaluation suite for intervention composability, introducing novel metrics including the *composability score*, and implement a modular PyTorch-based system with cross-category pipeline scheduling. Contribution/Results: Our analysis of 310 intervention combinations reveals critical interaction patterns: strong order dependence, suppression of editing and unlearning efficacy by compression, and the failure of conventional single-intervention metrics in compositional settings. All code is fully open-sourced, establishing a foundation for multi-objective, cooperative intervention paradigms in foundation models.

Develop framework for composable interventions with new metricsIdentify gaps in composability and need for multi-objective interventionsStudy interactions of multiple interventions on language models

Neuroplasticity and Corruption in Model Mechanisms: A Case Study Of Indirect Object Identification

Feb 27, 2025
VK
Vishnu Kabir Chhabra
🏛️ The Ohio State University

This study investigates the mechanistic degradation and reversibility of language models under toxic data fine-tuning. Toxic fine-tuning induces model corruption, yet its underlying neural mechanisms and potential for recovery remain poorly understood. Method: Leveraging causal tracing and circuit localization—key techniques from mechanistic interpretability—alongside task-specific fine-tuning and clean-data reverse retraining, we conduct controlled ablation and reconstruction experiments. Results: We establish, for the first time, that corruption exhibits *circuit-level specificity*: only critical computational pathways are selectively impaired, while peripheral circuits remain intact. Crucially, we demonstrate *neuroplastic-like recoverability*: clean-data retraining reconstructs original functional mechanisms with >89% restoration fidelity; this recovery generalizes across fine-tuning epochs. Contribution: Our work identifies precise circuit-level localization principles governing corruption and empirically validates the reversibility of mechanistic damage—providing both theoretical foundations and actionable strategies for robust alignment and trustworthy fine-tuning.

Identification of primary corruption mechanisms during toxic fine-tuning.Impact of fine-tuning on poisoned data and mechanism changes.Neuroplasticity behaviors in models retrained on clean datasets.

Towards Unifying Interpretability and Control: Evaluation via Intervention

Nov 07, 2024
UB
Usha Bhalla
🏛️ Harvard University | Google DeepMind

Current interpretability research for large language models (LLMs) treats interpretability and controllability as disjoint objectives. Method: This paper proposes “intervention capability” as a unified evaluation goal and introduces an encoder-decoder framework that integrates four method families—sparse autoencoders (SAEs), Logit Lens, Tuned Lens, and probes—to enable controllable interventions on interpretable features. Contribution/Results: We formally define two novel metrics—intervention success rate and consistency–intervention trade-off—and argue that effective intervention constitutes the foundational objective of interpretability. Experiments show that Lens-based methods outperform SAEs and probes in simple interventions; however, existing methods exhibit inconsistent cross-feature and cross-model intervention efficacy. Moreover, mechanistic interventions often underperform prompt engineering, revealing critical controllability bottlenecks. This work shifts LLM interpretability research from descriptive analysis toward causal, interventionist control.

Assess coherence-intervention tradeoff in model behaviorEvaluate methods through intervention success metricsUnify interpretability and control in language models

Latest Papers

What's happening recently
View more

We present an automated, contrastive evaluation pipeline for auditing the behavioral impact of interventions on large language models. Given a base model $M_1$ and an intervention model $M_2$, our method compares their free-form, multi-token generations across aligned prompt contexts and produces human-readable, statistically validated natural-language hypotheses describing how the models differ, along with recurring themes that summarize patterns across validated hypotheses. We evaluate the approach in synthetic setting by injecting known behavioral changes and showing that the pipeline reliably recovers them. We then apply it to three real-world interventions, reasoning distillation, knowledge editing and unlearning, demonstrating that the method surfaces both intended and unexpected behavioral shifts, distinguishes large from subtle interventions, and does not hallucinate differences when effects are absent or misaligned with the prompt bank. Overall, the pipeline provides a statistically grounded and interpretable tool for post-hoc auditing of intervention-induced changes in model behavior.

behavioral impactinterventionslanguage models

Activation patching is widely used in mechanistic interpretability, yet its natural indirect effect (NIE) estimator conflates inter-component state-dependent interaction effects (INT), leading to misattribution of causal contributions. This work formalizes and identifies INT within activation patching through the lens of causal mediation analysis, demonstrating that such interactions are both unavoidable and decomposable. The authors propose reframing INT as a diagnostic tool for interpretability. By integrating compositional interaction decomposition with local affine analyses, they empirically show in GPT-2’s IOI circuit that INT can cause critical components to be overlooked or their importance inflated, thereby explaining the instability of faithfulness scores. Building on these insights, they introduce new criteria for prompt-dependence analysis and mechanism discovery grounded in INT.

activation patchingcausal mediationinteraction effects

Existing intervention methods for language models predominantly rely on global linear directions, which fail to capture neuron-level nonlinear effects and cross-layer interactions, thereby limiting fine-grained control over model behavior. This work proposes Distributed Sparse Intervention (DSI), a method that performs sparse, structured causal interventions at the neuron level. DSI is the first to systematically model nonlinear interactions among neurons across layers and introduces set-theoretic operations to analyze the sparsity and composability of task representations. Experiments demonstrate that by intervening on only 0.01% of critical neurons, DSI can precisely activate desired behaviors across multiple tasks, significantly outperforming current linear intervention approaches while offering both high efficiency and strong interpretability.

distributed sparse interventionslanguage model steeringneuron-level interventions

Existing interpretability methods struggle to distinguish whether model components genuinely encode a target capability or merely propagate upstream signals. This work proposes Weight Patching, a source-directed intervention in weight space that operates on isomorphic models exhibiting varying behavioral strengths. By substituting specific module weights and anchoring behavioral interfaces via vector alignment, the method precisely localizes source-level mechanisms within large language models. The framework enables, for the first time, tracing the pathway of capability transmission from shallow source carriers to downstream execution circuits, thereby supporting mechanism-aware model merging. Experiments on instruction-following tasks successfully identify critical mechanistic components, significantly improving selective fusion of expert models, with findings further validated externally.

behavioral capabilityLLMsmechanistic interpretability

Hot Scholars

XW

Xiaozhi Wang

Tsinghua University
Natural Language ProcessingLanguage ModelMechanistic Interpretability
YB

Yonatan Belinkov

Technion
Natural Language ProcessingModel InterpretabilityArtificial Intelligence
RJ

Robin Jia

University of Southern California
natural language processing
JL

Juanzi Li

Tsinghua University
Semantic Webdata miningNLP