Score
Design, implement, and analyze targeted interventions on internal neural network activations—e.g., adding or replacing residual‑stream steering vectors, injecting attention‑logit biases, or editing representation readouts—to causally test how specific internal directions and components affect model outputs. Use these activation interventions both to steer or repair behavior (recovering readout‑correctable failures without retraining) and to quantify the causal field effects of representation changes.
This work proposes a practical, three-stage “Locate–Guide–Improve” framework that transforms mechanistic interpretability from a post-hoc diagnostic tool into an engineering-driven optimization methodology for large language models. By systematically integrating techniques for identifying critical neurons and pathways with targeted interventions—such as activation manipulation and module editing—the framework establishes a standardized protocol for model refinement while clearly distinguishing between localization and guidance mechanisms. Empirical results demonstrate significant improvements in model alignment, task performance, and reasoning efficiency, thereby advancing mechanistic interpretability toward real-world applicability.
This study investigates whether internal states of a language model, after activation intervention, can be reproduced through forward passes induced by arbitrary natural-language prompts. Formalizing this question as the surjectivity of prompt-induced state manifolds onto intervened states, we theoretically prove for the first time that activation interventions typically displace the residual stream outside the manifold of states reachable by any natural prompt, thereby rigorously delineating the capability boundaries between white-box interventions and black-box prompting. Leveraging the manifold hypothesis and probabilistic arguments, we empirically validate this theory across three prominent large language models, demonstrating that nearly all intervened states are irreproducible by any natural prompt. These findings indicate that successful activation interventions do not imply the existence of corresponding interpretable prompts or model vulnerabilities, revealing a fundamental separation between intervention and prompting mechanisms.
Existing activation intervention methods rely heavily on empirical design and lack theoretical grounding, making it difficult to systematically achieve efficient model adaptation. This work establishes, for the first time, a first-order equivalence between activation-space interventions and weight-space fine-tuning, revealing that the outputs of later transformer blocks serve as highly expressive intervention locations. Building on this insight, we propose a theoretically grounded strategy for selecting optimal intervention points. Furthermore, we introduce a novel weight-activation joint adaptation paradigm that simultaneously optimizes in both spaces. By training only 0.04% of the model parameters, our method achieves 99.1%–99.8% of the performance of full fine-tuning across multiple tasks, significantly outperforming mainstream parameter-efficient approaches such as ReFT and LoRA.
This work addresses the limited trustworthiness of Transformer models in high-stakes applications, which stems from insufficient understanding of their internal decision-making mechanisms. To bridge this gap, we propose a mechanistic interpretability approach based on targeted interventions on attention heads, integrating causal analysis with neural circuit probing to systematically uncover the model’s decision processes and underlying cognitive mechanisms. Our method substantially enhances the interpretability of Transformer internals and offers an innovative pathway toward the design and control of highly reliable AI systems, while also enabling the discovery of novel scientific insights encoded within these models.
This work addresses the limited capability of large language models (LLMs) to adhere to fine-grained instruction constraints—such as formatting, length, and keyword requirements—and their poor generalization across zero-shot or cross-model settings. To this end, we propose activation steering: a lightweight, inference-time intervention that computes layer-wise neural activation differences between instruction-present and instruction-absent conditions, yielding interpretable, transferable, and composable instruction vectors. Crucially, no model fine-tuning is required. Our key contribution is the first formulation of instructions as cross-model-transferable activation-difference vectors, enabling vector composition (e.g.,叠加 multiple constraints) and foundation-model enhancement. Extensive evaluation across four mainstream LLMs demonstrates substantial improvements in instruction-following accuracy. The method supports constraint-aware generation without explicit instructions, concurrent multi-constraint control, and knowledge transfer from instruction-tuned models to base models.
Current interpretability research for large language models (LLMs) treats interpretability and controllability as disjoint objectives. Method: This paper proposes “intervention capability” as a unified evaluation goal and introduces an encoder-decoder framework that integrates four method families—sparse autoencoders (SAEs), Logit Lens, Tuned Lens, and probes—to enable controllable interventions on interpretable features. Contribution/Results: We formally define two novel metrics—intervention success rate and consistency–intervention trade-off—and argue that effective intervention constitutes the foundational objective of interpretability. Experiments show that Lens-based methods outperform SAEs and probes in simple interventions; however, existing methods exhibit inconsistent cross-feature and cross-model intervention efficacy. Moreover, mechanistic interventions often underperform prompt engineering, revealing critical controllability bottlenecks. This work shifts LLM interpretability research from descriptive analysis toward causal, interventionist control.
This work addresses the prevalent issue in tool-augmented large language models (LLMs) of unnecessarily frequent external tool invocations even when tools are not required, reflecting a lack of precise control over tool-calling behavior. The authors propose a novel method that extracts activation steering vectors anchored at specific header positions in the input context. This approach reveals, for the first time, that although tool usage lacks explicit parametric encoding, it can be bidirectionally and causally modulated through activation vectors at context-dependent locations. Through activation steering, geometric representation analysis, and extensive cross-model and cross-domain experiments, the method significantly suppresses redundant tool calls across five open-source LLMs and three task domains. The findings demonstrate that internal representations governing tool use exhibit nonlinear, multimodal, and tool-type-specific characteristics, enabling precise and effective regulation of tool-invocation behavior.
Large language models may alter their behavior during safety evaluations due to awareness of being assessed, thereby compromising evaluation validity. This work proposes a novel method that suppresses internal latent variables associated with evaluation awareness solely by optimizing input prefix prompts, without requiring model inference access. The approach integrates GCG-style token optimization, a self-cross-entropy fluency regularizer, and multi-class latent targets—including CAA directions, SAE features, and MLP neurons—enabling, for the first time, selective deactivation of specific internal representations. Experiments on Llama-3.2-3B and Llama-3.1-8B demonstrate robust suppression of target latents to approximately –7, with causally validated SAE features fully deactivated, revealing that activation interpretability does not imply behavioral controllability.
This study addresses the lack of systematic analysis regarding the upstream sources of steering signals in activation intervention research, which has limited intervention efficacy. By fixing downstream intervention conditions and systematically manipulating source context and activation reading strategies, the work identifies the "execution boundary state" as a critical source of effective steering signals. To enhance signal purity and stability, the authors propose a tail-truncation method that disentangles prompt and continuation semantics. Experiments across three instruction-tuned models and four steering tasks demonstrate that judicious selection of source activations substantially improves intervention performance, with execution boundary states consistently outperforming contexts containing only target behaviors.