Score
Design and implement interventions that replace, edit, or transplant internal neural activations (e.g., overwriting with donor activations, mean vectors, or edited vectors, and performing back‑patching or single‑step interchange) and the tooling to apply those patches. Analyze how these activation-level interventions change model outputs to attribute behavior to components, estimate causal effects (including indirect effects), quantify intervention effectiveness across layers or authority levels, and test reversibility of erased representations.
Explanability research in large language models (LLMs) suffers from the absence of a unified theoretical framework, ambiguous definitions of causal mediators, and difficulties in cross-method comparison. Method: This work systematically introduces causal mediation analysis into LLM interpretability research—establishing, for the first time, a theoretical framework grounded in causal mediation. It proposes a two-dimensional taxonomy: one dimension classifies mediators by type (local vs. global, linear vs. nonlinear, explicit vs. implicit); the other categorizes search paradigms (gradient-based, perturbation-based, decomposition-based, activation-based). Contribution/Results: The framework introduces the novel “interpretability–efficiency trade-off” evaluation dimension, advocates standardized benchmarking, and advances modeling of higher-order, nonlinear abstract mediators. It provides principled guidance for method selection, clarifies the applicability boundaries of existing techniques, and charts a roadmap for next-generation interpretability methods grounded in causal reasoning.
This study addresses the challenge that internal interventions in visual world models struggle to sustain their effects during long-horizon autonomous prediction. To tackle this, it proposes the RolloutFaith framework, which introduces a novel metric for evaluating intervention persistence and benchmarks reference activation patching combined with low-rank editing in environments such as Crafter. Building upon these findings, the work further presents Delayed LoReFT, an optimized intervention method. The study reveals that current editors exhibit limited long-term efficacy, identifying newly generated frames and memory states as critical carriers of persistent effects. By validating the training hypothesis that rewarding future consequences enhances intervention durability, the proposed approach significantly improves persistence, thereby establishing a new paradigm for the reliable editing of world models.
Existing activation intervention methods rely heavily on empirical design and lack theoretical grounding, making it difficult to systematically achieve efficient model adaptation. This work establishes, for the first time, a first-order equivalence between activation-space interventions and weight-space fine-tuning, revealing that the outputs of later transformer blocks serve as highly expressive intervention locations. Building on this insight, we propose a theoretically grounded strategy for selecting optimal intervention points. Furthermore, we introduce a novel weight-activation joint adaptation paradigm that simultaneously optimizes in both spaces. By training only 0.04% of the model parameters, our method achieves 99.1%–99.8% of the performance of full fine-tuning across multiple tasks, significantly outperforming mainstream parameter-efficient approaches such as ReFT and LoRA.
本文通过因果中介分析方法,诊断视觉-语言模型中的性别偏见机制,揭示了各层激活对干预的敏感性。
This study addresses the latent degradation and evaluation blind spots induced by internal activation steering in tool-calling scenarios of large language models (LLMs). We propose SAKIKO, an auditing framework that systematically evaluates the genuine remediation effects of internal interventions through directional error discovery, channel-keyed intervention, and target-parsing verification. Furthermore, this work pioneers an outcome-parsing-based adjudication mechanism and forward-freezing statistical licensing, revealing that behavioral modification does not equate to fundamental repair. Experiments across seven LLMs demonstrate that most existing interventions incur severe side effects, thereby establishing the necessity of outcome-level adjudication for ensuring intervention rigor.
Existing test-time interventions—such as knowledge editing, model compression, and machine unlearning—have evolved in isolation, with no standardized framework to systematically study their interactions when applied jointly on the same model. Method: We propose the first composable intervention framework, unifying the modeling of synergistic mechanisms across these three intervention types. We design a comprehensive evaluation suite for intervention composability, introducing novel metrics including the *composability score*, and implement a modular PyTorch-based system with cross-category pipeline scheduling. Contribution/Results: Our analysis of 310 intervention combinations reveals critical interaction patterns: strong order dependence, suppression of editing and unlearning efficacy by compression, and the failure of conventional single-intervention metrics in compositional settings. All code is fully open-sourced, establishing a foundation for multi-objective, cooperative intervention paradigms in foundation models.
We present an automated, contrastive evaluation pipeline for auditing the behavioral impact of interventions on large language models. Given a base model $M_1$ and an intervention model $M_2$, our method compares their free-form, multi-token generations across aligned prompt contexts and produces human-readable, statistically validated natural-language hypotheses describing how the models differ, along with recurring themes that summarize patterns across validated hypotheses. We evaluate the approach in synthetic setting by injecting known behavioral changes and showing that the pipeline reliably recovers them. We then apply it to three real-world interventions, reasoning distillation, knowledge editing and unlearning, demonstrating that the method surfaces both intended and unexpected behavioral shifts, distinguishes large from subtle interventions, and does not hallucinate differences when effects are absent or misaligned with the prompt bank. Overall, the pipeline provides a statistically grounded and interpretable tool for post-hoc auditing of intervention-induced changes in model behavior.
本文提出了一种基于AI的神经替代框架,通过fMRI活动快照预测感知效应来设计认知-情感神经调节目标,无需物理刺激。
Activation patching is widely used in mechanistic interpretability, yet its natural indirect effect (NIE) estimator conflates inter-component state-dependent interaction effects (INT), leading to misattribution of causal contributions. This work formalizes and identifies INT within activation patching through the lens of causal mediation analysis, demonstrating that such interactions are both unavoidable and decomposable. The authors propose reframing INT as a diagnostic tool for interpretability. By integrating compositional interaction decomposition with local affine analyses, they empirically show in GPT-2’s IOI circuit that INT can cause critical components to be overlooked or their importance inflated, thereby explaining the instability of faithfulness scores. Building on these insights, they introduce new criteria for prompt-dependence analysis and mechanism discovery grounded in INT.
This study addresses the challenge in large language model (LLM) agents of distinguishing instructions from data and localizing intervention points under indirect prompt injection. By employing counterfactual role probes, component-level activation patching, and trajectory-independent interventions via AgentDojo, this work leverages causal tracing to reveal a fundamental divergence between the readability of role signals and the intervenability of agent behavior. It is the first to explicitly distinguish readable role signals from effective behavioral interventions, demonstrating that cross-channel transfer of intervention directions is inherently difficult. Furthermore, the study systematically analyzes how network depth and positional factors influence attack success. Results indicate that single-point editing fails in deeper layers, whereas wide-span and repeated editing strategies significantly reduce attack success rates, offering actionable insights for securing LLM agents against indirect prompt injection.
This study addresses the lack of functional self-awareness in neural networks by proposing a self-intervention learning framework. Introducing a novel autonomous perturbation mechanism, the model constructs a predictive self-model to facilitate introspective learning and guide structural optimization. This approach integrates causal perturbation with self-modeling, establishing a new paradigm that derives self-knowledge from functional consequences. Experimental results demonstrate that the self-model significantly reduces prediction error, decreasing action regret by 31.7%. Although it does not comprehensively outperform direct policy methods, this work successfully validates the effectiveness of structured adjustments based on predictive self-awareness, offering a promising avenue for research into neural network introspection capabilities.