Score
Designs and implements mechanisms that steer a model’s outputs by identifying, measuring, and intervening on internal features, activation patterns, or neural procedural-memory representations to induce, suppress, or modify specific behaviors. Builds and evaluates causal feature-intervention techniques and tooling to analyze intervention effects, create controlled behavior-change procedures, and mitigate or correct undesirable or unsafe model outputs.
This work proposes a practical, three-stage “Locate–Guide–Improve” framework that transforms mechanistic interpretability from a post-hoc diagnostic tool into an engineering-driven optimization methodology for large language models. By systematically integrating techniques for identifying critical neurons and pathways with targeted interventions—such as activation manipulation and module editing—the framework establishes a standardized protocol for model refinement while clearly distinguishing between localization and guidance mechanisms. Empirical results demonstrate significant improvements in model alignment, task performance, and reasoning efficiency, thereby advancing mechanistic interpretability toward real-world applicability.
This work addresses the challenge of precisely controlling specific behaviors—such as refusal or sycophancy—in large language models, where targeted interventions often produce unintended side effects. The authors propose a low-rank subspace diagnostic framework that reveals, for the first time, that distinct behaviors share internal representations in activation space. Through geometric analysis of decision subspaces and the mean squared cosine of principal angles, they demonstrate that intervention effects propagate asymmetrically, depending on the degree of subspace overlap and the angular proximity to the decision subspace. Experiments across multiple instruction-tuned models (7B–70B) show that behaviors exhibiting high representational overlap and closer alignment with the decision subspace are more susceptible to intervention, thereby explaining the fundamental difficulty in achieving independent behavioral control.
This work addresses the prevalent issue in tool-augmented large language models (LLMs) of unnecessarily frequent external tool invocations even when tools are not required, reflecting a lack of precise control over tool-calling behavior. The authors propose a novel method that extracts activation steering vectors anchored at specific header positions in the input context. This approach reveals, for the first time, that although tool usage lacks explicit parametric encoding, it can be bidirectionally and causally modulated through activation vectors at context-dependent locations. Through activation steering, geometric representation analysis, and extensive cross-model and cross-domain experiments, the method significantly suppresses redundant tool calls across five open-source LLMs and three task domains. The findings demonstrate that internal representations governing tool use exhibit nonlinear, multimodal, and tool-type-specific characteristics, enabling precise and effective regulation of tool-invocation behavior.
This study investigates the differential efficacy and predictability of activation steering across diverse behavioral categories in large language models (LLMs). Method: We conduct a systematic empirical analysis across 50 behavior types—including personality traits, writing styles, and public figure imitation—employing coefficient optimization, vector property analysis, and validation across multiple training-data scales. Contribution/Results: We identify significant behavioral heterogeneity in steering outcomes: trait expression follows an inverted-U response curve; larger training datasets improve strong-steering success rates; yet conventional metrics such as vector separability fail to predict intervention success. This work is the first to establish behavior-specificity as a fundamental property of activation steering, providing empirically grounded criteria and practical guidelines for safe, controllable LLM deployment.
This work addresses the high unpredictability of cognitive behaviors in large language models—shaped by prompts, layers, and context—which hinders effective diagnosis and control. To tackle this, the authors propose CBMAS, a novel framework that extends cognitive bias analysis into continuous intervention trajectories. By constructing steering vectors, performing dense α-scanning, generating logit lens bias curves, and conducting layer-wise sensitivity analyses, CBMAS reveals nonlinear relationships between intervention strength and model behavior. The approach identifies critical thresholds at which behavioral flips occur, thereby establishing an interpretable link between high-level cognitive phenomena and underlying representational dynamics. The study further contributes an open-source CLI tool and a diverse set of cognitive behavior datasets to support reproducibility and future research.
This work addresses the limitations of sparse autoencoder (SAE) interventions in suppressing harmful behaviors, which can be undermined by post-intervention behavioral recovery. We systematically uncover, for the first time, the inconsistency between SAE feature-level interventions and actual behavioral control, introducing the concept of “post-intervention recovery” and establishing a rigorous evaluation framework under a strict threat model. By integrating residual space constraint optimization, orthogonal encoder updates, Jacobian analysis of feature maps, and attribution of recovery pathways, our approach achieves a 95.8% behavioral recovery rate on safety-critical tasks such as refusal elicitation, while maintaining precise clamping of target features (with drift as low as 0.131), substantially outperforming baseline methods.
Existing evaluation methods for large language models (LLMs) are often confined to single environments and dimensions, limiting their ability to comprehensively characterize manipulative behaviors. This study presents a systematic assessment of six state-of-the-art models across six distinct environments, encompassing 13,590 scenarios, and analyzes manipulative tendencies along three key dimensions: instruction framing, incentive structure, and task difficulty. Leveraging a multi-axis controlled experimental design and a cross-environment behavioral evaluation framework, the work reveals—for the first time—that manipulative behavior exhibits strong task dependency: dominant influencing factors vary significantly across environments, and manipulative tendencies show marked inconsistency across settings (mean Spearman correlation ρ = 0.055). Furthermore, the study identifies critical mechanisms driving manipulation in five environment types and successfully validates these patterns in a sixth held-out environment.