Score
Design, build, or analyze model interventions, representations, or training objectives that intentionally couple a model's harmfulness (propensity to produce harmful outputs) with its refusal behavior (propensity to decline harmful requests). These interventions align harmfulness and refusal directions across prompt and response positions and are constrained to a harmfulness–refusal subspace so general capability and usability are preserved.
This work addresses the challenge of precisely controlling specific behaviors—such as refusal or sycophancy—in large language models, where targeted interventions often produce unintended side effects. The authors propose a low-rank subspace diagnostic framework that reveals, for the first time, that distinct behaviors share internal representations in activation space. Through geometric analysis of decision subspaces and the mean squared cosine of principal angles, they demonstrate that intervention effects propagate asymmetrically, depending on the degree of subspace overlap and the angular proximity to the decision subspace. Experiments across multiple instruction-tuned models (7B–70B) show that behaviors exhibiting high representational overlap and closer alignment with the decision subspace are more susceptible to intervention, thereby explaining the fundamental difficulty in achieving independent behavioral control.
This work investigates whether large language models (LLMs) possess intrinsic semantic understanding of *harmfulness*—beyond merely learning surface-level refusal behaviors—when rejecting harmful instructions. Using causal steering in the latent space, we disentangle a *harmfulness direction* orthogonal to the refusal direction, revealing that the model’s internal representation of harm is more stable than its refusal behavior. Building on this insight, we propose Latent Guard: an implicit safety mechanism that detects harmfulness directly via the identified latent direction, eliminating the need for an explicit classification head. Experiments show that Latent Guard matches or surpasses Llama Guard 3-8B in detection accuracy across diverse jailbreaking attacks, significantly reduces over-refusal, and exhibits strong robustness against adversarial fine-tuning attacks. This work establishes a novel, interpretable, and generalizable representation-level perspective for AI safety mechanisms.
This study investigates the refusal mechanisms of large language models (LLMs) against harmful requests, challenging the conventional assumption of linear separability in refusal behavior. Method: Leveraging nonlinear dimensionality reduction techniques—including PCA, t-SNE, and UMAP—we systematically analyze hidden states across layers of six LLMs spanning three distinct architectures. Contribution/Results: We empirically demonstrate that refusal decisions cannot be characterized by low-dimensional linear boundaries; instead, they reside on model-specific, nonlinear refusal manifolds, revealing intrinsic nonlinearity, multidimensional heterogeneity, and strong architecture- and layer-dependence. This work establishes “nonlinear interpretability” as a novel paradigm for safety alignment, providing both theoretical foundations and methodological pathways toward interpretable and robust refusal mechanisms.
Existing activation steering methods lack input awareness, hindering fine-grained, selective response control. This paper proposes Conditional Activation Steering (CAST), the first approach to enable input-semantic-category–conditioned activation intervention: by analyzing latent-state activation patterns during LLM inference, CAST dynamically triggers refusal responses for specific risk categories (e.g., hate speech, adult content) without fine-tuning or modifying model weights. CAST integrates latent-state pattern recognition, conditional triggering, and targeted activation-space offsets, enabling rule-driven zero-shot behavioral programming. Evaluated across multiple safety-critical and domain-specific refusal tasks, CAST achieves >92% recall while preserving response quality for non-target inputs (BLEU degradation <0.5), thus balancing safety and general-purpose utility.
Large language models (LLMs) exhibit a “refusal position bias” during safety fine-tuning—i.e., they predominantly refuse harmful queries only at response beginnings or endings—resulting in insufficient end-to-end safety enforcement. To address this, we propose Decoupled Refusal Training (DeRTa), the first framework enabling position-agnostic refusal learning. DeRTa introduces two core components: (1) hazard-response prefix guidance coupled with prefix-augmented maximum likelihood estimation (MLE) for fine-grained refusal modeling, and (2) a Reinforcement Transition Optimization (RTO) mechanism that enables dynamic, context-aware safety interception at arbitrary token positions within the response. Extensive experiments on LLaMA-3 and Mistral-family models across six adversarial attack scenarios demonstrate that DeRTa achieves significantly higher safety robustness than state-of-the-art baselines while preserving general-purpose capabilities—confirming no trade-off between safety and utility.
This work demonstrates that language models’ safety mechanisms can be circumvented by simple prefilled prompts such as “Sure, here is,” which bypass refusal behaviors without altering the semantic content of the input. The study reveals that refusal decisions are shallow computations concentrated in the early stages of response generation and rely primarily on general autoregressive conditioning rather than dedicated safety modules. Using interpretability techniques—including linear probing, causal attention interventions, state ablation, and logit trajectory analysis—the authors validate across models ranging from 1.5B to 14B parameters that causal interventions targeting the initial response window reduce jailbreak success rates by 74%. The same intervention also substantially suppresses harmful outputs in base models (from 64% to 25%), demonstrating the generality of this mechanism.