harmfulness-refusal coupling

Design, build, or analyze model interventions, representations, or training objectives that intentionally couple a model's harmfulness (propensity to produce harmful outputs) with its refusal behavior (propensity to decline harmful requests). These interventions align harmfulness and refusal directions across prompt and response positions and are constrained to a harmfulness–refusal subspace so general capability and usability are preserved.

harmfulness-refusalcoupling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of precisely controlling specific behaviors—such as refusal or sycophancy—in large language models, where targeted interventions often produce unintended side effects. The authors propose a low-rank subspace diagnostic framework that reveals, for the first time, that distinct behaviors share internal representations in activation space. Through geometric analysis of decision subspaces and the mean squared cosine of principal angles, they demonstrate that intervention effects propagate asymmetrically, depending on the degree of subspace overlap and the angular proximity to the decision subspace. Experiments across multiple instruction-tuned models (7B–70B) show that behaviors exhibiting high representational overlap and closer alignment with the decision subspace are more susceptible to intervention, thereby explaining the fundamental difficulty in achieving independent behavioral control.

behavioral interferencelarge language modelssafety interventions

LLMs Encode Harmfulness and Refusal Separately

Jul 15, 2025
JZ
Jiachen Zhao
🏛️ Northeastern University | Stanford University

This work investigates whether large language models (LLMs) possess intrinsic semantic understanding of *harmfulness*—beyond merely learning surface-level refusal behaviors—when rejecting harmful instructions. Using causal steering in the latent space, we disentangle a *harmfulness direction* orthogonal to the refusal direction, revealing that the model’s internal representation of harm is more stable than its refusal behavior. Building on this insight, we propose Latent Guard: an implicit safety mechanism that detects harmfulness directly via the identified latent direction, eliminating the need for an explicit classification head. Experiments show that Latent Guard matches or surpasses Llama Guard 3-8B in detection accuracy across diverse jailbreaking attacks, significantly reduces over-refusal, and exhibits strong robustness against adversarial fine-tuning attacks. This work establishes a novel, interpretable, and generalizable representation-level perspective for AI safety mechanisms.

Analyze how jailbreak methods bypass refusal signalsIdentify harmfulness as separate from refusal in LLMsPropose latent harmfulness representation for robust safety

Refusal Behavior in Large Language Models: A Nonlinear Perspective

Jan 14, 2025
FH
Fabian Hildebrandt
🏛️ FAU Erlangen-Nürnberg | University Hospital Erlangen

This study investigates the refusal mechanisms of large language models (LLMs) against harmful requests, challenging the conventional assumption of linear separability in refusal behavior. Method: Leveraging nonlinear dimensionality reduction techniques—including PCA, t-SNE, and UMAP—we systematically analyze hidden states across layers of six LLMs spanning three distinct architectures. Contribution/Results: We empirically demonstrate that refusal decisions cannot be characterized by low-dimensional linear boundaries; instead, they reside on model-specific, nonlinear refusal manifolds, revealing intrinsic nonlinearity, multidimensional heterogeneity, and strong architecture- and layer-dependence. This work establishes “nonlinear interpretability” as a novel paradigm for safety alignment, providing both theoretical foundations and methodological pathways toward interpretable and robust refusal mechanisms.

Decision MechanismLarge Language ModelsSafety

Programming Refusal with Conditional Activation Steering

Sep 06, 2024
BW
Bruce W. Lee
🏛️ University of Pennsylvania | IBM Research

Existing activation steering methods lack input awareness, hindering fine-grained, selective response control. This paper proposes Conditional Activation Steering (CAST), the first approach to enable input-semantic-category–conditioned activation intervention: by analyzing latent-state activation patterns during LLM inference, CAST dynamically triggers refusal responses for specific risk categories (e.g., hate speech, adult content) without fine-tuning or modifying model weights. CAST integrates latent-state pattern recognition, conditional triggering, and targeted activation-space offsets, enabling rule-driven zero-shot behavioral programming. Evaluated across multiple safety-critical and domain-specific refusal tasks, CAST achieves >92% recall while preserving response quality for non-target inputs (BLEU degradation <0.5), thus balancing safety and general-purpose utility.

Conditional Activation Steering methodDomain-specific behavior modificationSelective response control in LLMs

Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training

Jul 12, 2024
YY
Youliang Yuan
🏛️ The Chinese University of Hong Kong | Tencent AI Lab | Shenzhen Research Institute of Big Data

Large language models (LLMs) exhibit a “refusal position bias” during safety fine-tuning—i.e., they predominantly refuse harmful queries only at response beginnings or endings—resulting in insufficient end-to-end safety enforcement. To address this, we propose Decoupled Refusal Training (DeRTa), the first framework enabling position-agnostic refusal learning. DeRTa introduces two core components: (1) hazard-response prefix guidance coupled with prefix-augmented maximum likelihood estimation (MLE) for fine-grained refusal modeling, and (2) a Reinforcement Transition Optimization (RTO) mechanism that enables dynamic, context-aware safety interception at arbitrary token positions within the response. Extensive experiments on LLaMA-3 and Mistral-family models across six adversarial attack scenarios demonstrate that DeRTa achieves significantly higher safety robustness than state-of-the-art baselines while preserving general-purpose capabilities—confirming no trade-off between safety and utility.

Addresses refusal position bias in LLM safety tuningEnhances LLM ability to refuse harmful prompts flexiblyImproves safety without compromising model performance

Latest Papers

What's happening recently
View more

This work demonstrates that language models’ safety mechanisms can be circumvented by simple prefilled prompts such as “Sure, here is,” which bypass refusal behaviors without altering the semantic content of the input. The study reveals that refusal decisions are shallow computations concentrated in the early stages of response generation and rely primarily on general autoregressive conditioning rather than dedicated safety modules. Using interpretability techniques—including linear probing, causal attention interventions, state ablation, and logit trajectory analysis—the authors validate across models ranging from 1.5B to 14B parameters that causal interventions targeting the initial response window reduce jailbreak success rates by 74%. The same intervention also substantially suppresses harmful outputs in base models (from 64% to 25%), demonstrating the generality of this mechanism.

aligned language modelsharm representationjailbreak