Score
Design and implement interventions that replace, edit, or transplant internal neural activations (e.g., overwriting with donor activations, mean vectors, or edited vectors, and performing back‑patching or single‑step interchange) and the tooling to apply those patches. Analyze how these activation-level interventions change model outputs to attribute behavior to components, estimate causal effects (including indirect effects), quantify intervention effectiveness across layers or authority levels, and test reversibility of erased representations.
Existing activation intervention methods rely heavily on empirical design and lack theoretical grounding, making it difficult to systematically achieve efficient model adaptation. This work establishes, for the first time, a first-order equivalence between activation-space interventions and weight-space fine-tuning, revealing that the outputs of later transformer blocks serve as highly expressive intervention locations. Building on this insight, we propose a theoretically grounded strategy for selecting optimal intervention points. Furthermore, we introduce a novel weight-activation joint adaptation paradigm that simultaneously optimizes in both spaces. By training only 0.04% of the model parameters, our method achieves 99.1%–99.8% of the performance of full fine-tuning across multiple tasks, significantly outperforming mainstream parameter-efficient approaches such as ReFT and LoRA.
Existing test-time interventions—such as knowledge editing, model compression, and machine unlearning—have evolved in isolation, with no standardized framework to systematically study their interactions when applied jointly on the same model. Method: We propose the first composable intervention framework, unifying the modeling of synergistic mechanisms across these three intervention types. We design a comprehensive evaluation suite for intervention composability, introducing novel metrics including the *composability score*, and implement a modular PyTorch-based system with cross-category pipeline scheduling. Contribution/Results: Our analysis of 310 intervention combinations reveals critical interaction patterns: strong order dependence, suppression of editing and unlearning efficacy by compression, and the failure of conventional single-intervention metrics in compositional settings. All code is fully open-sourced, establishing a foundation for multi-objective, cooperative intervention paradigms in foundation models.
This study investigates the mechanistic degradation and reversibility of language models under toxic data fine-tuning. Toxic fine-tuning induces model corruption, yet its underlying neural mechanisms and potential for recovery remain poorly understood. Method: Leveraging causal tracing and circuit localization—key techniques from mechanistic interpretability—alongside task-specific fine-tuning and clean-data reverse retraining, we conduct controlled ablation and reconstruction experiments. Results: We establish, for the first time, that corruption exhibits *circuit-level specificity*: only critical computational pathways are selectively impaired, while peripheral circuits remain intact. Crucially, we demonstrate *neuroplastic-like recoverability*: clean-data retraining reconstructs original functional mechanisms with >89% restoration fidelity; this recovery generalizes across fine-tuning epochs. Contribution: Our work identifies precise circuit-level localization principles governing corruption and empirically validates the reversibility of mechanistic damage—providing both theoretical foundations and actionable strategies for robust alignment and trustworthy fine-tuning.
Current interpretability research for large language models (LLMs) treats interpretability and controllability as disjoint objectives. Method: This paper proposes “intervention capability” as a unified evaluation goal and introduces an encoder-decoder framework that integrates four method families—sparse autoencoders (SAEs), Logit Lens, Tuned Lens, and probes—to enable controllable interventions on interpretable features. Contribution/Results: We formally define two novel metrics—intervention success rate and consistency–intervention trade-off—and argue that effective intervention constitutes the foundational objective of interpretability. Experiments show that Lens-based methods outperform SAEs and probes in simple interventions; however, existing methods exhibit inconsistent cross-feature and cross-model intervention efficacy. Moreover, mechanistic interventions often underperform prompt engineering, revealing critical controllability bottlenecks. This work shifts LLM interpretability research from descriptive analysis toward causal, interventionist control.
We present an automated, contrastive evaluation pipeline for auditing the behavioral impact of interventions on large language models. Given a base model $M_1$ and an intervention model $M_2$, our method compares their free-form, multi-token generations across aligned prompt contexts and produces human-readable, statistically validated natural-language hypotheses describing how the models differ, along with recurring themes that summarize patterns across validated hypotheses. We evaluate the approach in synthetic setting by injecting known behavioral changes and showing that the pipeline reliably recovers them. We then apply it to three real-world interventions, reasoning distillation, knowledge editing and unlearning, demonstrating that the method surfaces both intended and unexpected behavioral shifts, distinguishes large from subtle interventions, and does not hallucinate differences when effects are absent or misaligned with the prompt bank. Overall, the pipeline provides a statistically grounded and interpretable tool for post-hoc auditing of intervention-induced changes in model behavior.
Activation patching is widely used in mechanistic interpretability, yet its natural indirect effect (NIE) estimator conflates inter-component state-dependent interaction effects (INT), leading to misattribution of causal contributions. This work formalizes and identifies INT within activation patching through the lens of causal mediation analysis, demonstrating that such interactions are both unavoidable and decomposable. The authors propose reframing INT as a diagnostic tool for interpretability. By integrating compositional interaction decomposition with local affine analyses, they empirically show in GPT-2’s IOI circuit that INT can cause critical components to be overlooked or their importance inflated, thereby explaining the instability of faithfulness scores. Building on these insights, they introduce new criteria for prompt-dependence analysis and mechanism discovery grounded in INT.
Existing intervention methods for language models predominantly rely on global linear directions, which fail to capture neuron-level nonlinear effects and cross-layer interactions, thereby limiting fine-grained control over model behavior. This work proposes Distributed Sparse Intervention (DSI), a method that performs sparse, structured causal interventions at the neuron level. DSI is the first to systematically model nonlinear interactions among neurons across layers and introduces set-theoretic operations to analyze the sparsity and composability of task representations. Experiments demonstrate that by intervening on only 0.01% of critical neurons, DSI can precisely activate desired behaviors across multiple tasks, significantly outperforming current linear intervention approaches while offering both high efficiency and strong interpretability.
Existing interpretability methods struggle to distinguish whether model components genuinely encode a target capability or merely propagate upstream signals. This work proposes Weight Patching, a source-directed intervention in weight space that operates on isomorphic models exhibiting varying behavioral strengths. By substituting specific module weights and anchoring behavioral interfaces via vector alignment, the method precisely localizes source-level mechanisms within large language models. The framework enables, for the first time, tracing the pathway of capability transmission from shallow source carriers to downstream execution circuits, thereby supporting mechanism-aware model merging. Experiments on instruction-following tasks successfully identify critical mechanistic components, significantly improving selective fusion of expert models, with findings further validated externally.