Score
Design, build, and analyze interventions that compute, extract, and apply direction vectors or attention-map modifications to internal neural activations (e.g., concept activation vectors, cross‑attention steering, latent/activation vectors) so as to steer model behavior at inference without finetuning. Implement and evaluate methods for injecting, interpolating, orthogonalizing, sparsifying, or otherwise constraining additive or suppressive activation updates (including bidirectional and gated schemes), control steering strength and thresholds, minimize per‑token perturbation via closed‑form or constrained updates, and measure the causal effects of those interventions on downstream outputs (including using targeted synthetic examples for training or evaluation).
This work proposes a practical, three-stage “Locate–Guide–Improve” framework that transforms mechanistic interpretability from a post-hoc diagnostic tool into an engineering-driven optimization methodology for large language models. By systematically integrating techniques for identifying critical neurons and pathways with targeted interventions—such as activation manipulation and module editing—the framework establishes a standardized protocol for model refinement while clearly distinguishing between localization and guidance mechanisms. Empirical results demonstrate significant improvements in model alignment, task performance, and reasoning efficiency, thereby advancing mechanistic interpretability toward real-world applicability.
Existing activation intervention methods are constrained by oversimplified assumptions—namely, fixed, single-step, and position-invariant transformations—resulting in limited generalization and inferior performance compared to in-context prompting. This work proposes FLAS, the first approach to incorporate continuous flow fields into activation intervention. FLAS models the transformation from original to target activations via concept-conditional ordinary differential equations, learning a nonlinear, multi-step, token-adaptive mapping without requiring parameter freezing or per-concept fine-tuning. Evaluated on AxBench, FLAS achieves state-of-the-art results, surpassing in-context prompting for the first time, with harmonic mean scores of 1.015 and 1.113 on the Gemma-2-2B-IT and 9B-IT models, respectively.
Existing activation intervention methods rely heavily on empirical design and lack theoretical grounding, making it difficult to systematically achieve efficient model adaptation. This work establishes, for the first time, a first-order equivalence between activation-space interventions and weight-space fine-tuning, revealing that the outputs of later transformer blocks serve as highly expressive intervention locations. Building on this insight, we propose a theoretically grounded strategy for selecting optimal intervention points. Furthermore, we introduce a novel weight-activation joint adaptation paradigm that simultaneously optimizes in both spaces. By training only 0.04% of the model parameters, our method achieves 99.1%–99.8% of the performance of full fine-tuning across multiple tasks, significantly outperforming mainstream parameter-efficient approaches such as ReFT and LoRA.
To address the challenge of jointly achieving controllability and efficiency in generative models, this paper proposes LinEAS: a framework for sparse, end-to-end learnable linear interventions applied directly at activation layers. Its core innovation lies in a global loss-driven mechanism for automatic neuron- and layer-level selection, integrated with inter-layer distribution alignment loss, L1/L0 sparsity regularization, and few-shot activation-space calibration. LinEAS supports intervention composition and cross-task transfer, significantly enhancing robustness and overcoming the limitations of conventional local interventions. Experiments demonstrate that, using only a few samples, LinEAS substantially outperforms existing intervention methods on text detoxification and text-to-image style control—achieving toxicity reduction on par with full-parameter fine-tuning. Moreover, it successfully transfers to diffusion models such as Stable Diffusion, delivering both low computational overhead and high generation quality.
This work addresses the limited capability of large language models (LLMs) to adhere to fine-grained instruction constraints—such as formatting, length, and keyword requirements—and their poor generalization across zero-shot or cross-model settings. To this end, we propose activation steering: a lightweight, inference-time intervention that computes layer-wise neural activation differences between instruction-present and instruction-absent conditions, yielding interpretable, transferable, and composable instruction vectors. Crucially, no model fine-tuning is required. Our key contribution is the first formulation of instructions as cross-model-transferable activation-difference vectors, enabling vector composition (e.g.,叠加 multiple constraints) and foundation-model enhancement. Extensive evaluation across four mainstream LLMs demonstrates substantial improvements in instruction-following accuracy. The method supports constraint-aware generation without explicit instructions, concurrent multi-constraint control, and knowledge transfer from instruction-tuned models to base models.
Existing activation steering methods rely on globally fixed interventions, which lack fine-grained control and often compromise model utility. This work proposes Steer2Edit, a novel framework that establishes, for the first time, a theoretical connection between activation steering and weight editing. By reinterpreting steering vectors as diagnostic signals, Steer2Edit enables component-level rank-1 weight edits to attention heads and MLP neurons without any training, achieving localized and interpretable behavioral control while preserving standard forward propagation and inference efficiency. Experiments demonstrate that Steer2Edit significantly outperforms baselines across safety alignment, hallucination mitigation, and inference efficiency tasks—improving safety by up to 17.2%, factual accuracy by 9.8%, and reducing average inference length by 12.2%, all without degrading downstream performance.
Existing activation intervention methods struggle to disentangle angular and norm components in concept representations, leading to unstable control effects and opaque mechanisms. This work proposes an angle-norm decomposition framework and, through controlled experiments and evaluations across seven language models, empirically demonstrates that conceptual information is primarily encoded in the angular structure, while the norm plays a critical role in intervention stability. The study elucidates the geometric origins underlying the performance differences among various intervention approaches and advocates for parameterizing interventions via interpretable angular and radial components. This provides both theoretical grounding and practical guidance for achieving more stable and interpretable control over language model behaviors.
This work addresses the prevalent issue in tool-augmented large language models (LLMs) of unnecessarily frequent external tool invocations even when tools are not required, reflecting a lack of precise control over tool-calling behavior. The authors propose a novel method that extracts activation steering vectors anchored at specific header positions in the input context. This approach reveals, for the first time, that although tool usage lacks explicit parametric encoding, it can be bidirectionally and causally modulated through activation vectors at context-dependent locations. Through activation steering, geometric representation analysis, and extensive cross-model and cross-domain experiments, the method significantly suppresses redundant tool calls across five open-source LLMs and three task domains. The findings demonstrate that internal representations governing tool use exhibit nonlinear, multimodal, and tool-type-specific characteristics, enabling precise and effective regulation of tool-invocation behavior.
This study addresses the lack of systematic analysis regarding the upstream sources of steering signals in activation intervention research, which has limited intervention efficacy. By fixing downstream intervention conditions and systematically manipulating source context and activation reading strategies, the work identifies the "execution boundary state" as a critical source of effective steering signals. To enhance signal purity and stability, the authors propose a tail-truncation method that disentangles prompt and continuation semantics. Experiments across three instruction-tuned models and four steering tasks demonstrate that judicious selection of source activations substantially improves intervention performance, with execution boundary states consistently outperforming contexts containing only target behaviors.