refusal-focused fine-tuning

Designs, implements, and evaluates fine-tuning procedures, datasets, label schemes, loss functions, and training schedules that bias a model toward refusing or referring disallowed, harmful, or sensitive prompts; this work includes specifying refusal/referral signals, measuring refusal and referral rates and reductions in inappropriate response relevance, and optimizing behavior under compute or resource constraints.

refusal-focusedfine-tuning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Programming Refusal with Conditional Activation Steering

Sep 06, 2024
BW
Bruce W. Lee
🏛️ University of Pennsylvania | IBM Research

Existing activation steering methods lack input awareness, hindering fine-grained, selective response control. This paper proposes Conditional Activation Steering (CAST), the first approach to enable input-semantic-category–conditioned activation intervention: by analyzing latent-state activation patterns during LLM inference, CAST dynamically triggers refusal responses for specific risk categories (e.g., hate speech, adult content) without fine-tuning or modifying model weights. CAST integrates latent-state pattern recognition, conditional triggering, and targeted activation-space offsets, enabling rule-driven zero-shot behavioral programming. Evaluated across multiple safety-critical and domain-specific refusal tasks, CAST achieves >92% recall while preserving response quality for non-target inputs (BLEU degradation <0.5), thus balancing safety and general-purpose utility.

Conditional Activation Steering methodDomain-specific behavior modificationSelective response control in LLMs

Refusal Behavior in Large Language Models: A Nonlinear Perspective

Jan 14, 2025
FH
Fabian Hildebrandt
🏛️ FAU Erlangen-Nürnberg | University Hospital Erlangen

This study investigates the refusal mechanisms of large language models (LLMs) against harmful requests, challenging the conventional assumption of linear separability in refusal behavior. Method: Leveraging nonlinear dimensionality reduction techniques—including PCA, t-SNE, and UMAP—we systematically analyze hidden states across layers of six LLMs spanning three distinct architectures. Contribution/Results: We empirically demonstrate that refusal decisions cannot be characterized by low-dimensional linear boundaries; instead, they reside on model-specific, nonlinear refusal manifolds, revealing intrinsic nonlinearity, multidimensional heterogeneity, and strong architecture- and layer-dependence. This work establishes “nonlinear interpretability” as a novel paradigm for safety alignment, providing both theoretical foundations and methodological pathways toward interpretable and robust refusal mechanisms.

Decision MechanismLarge Language ModelsSafety

Characterizing Selective Refusal Bias in Large Language Models

Oct 30, 2025
AK
Adel Khorramrouz
🏛️ Rutgers University

This work identifies selective refusal bias in large language model (LLM) safety mechanisms: systematic disparities exist across demographic groups—by gender, sexual orientation, nationality, and religion—in refusal rates, response types, and refusal text length when generating harmful content, resulting in inadequate protection for marginalized populations. We introduce the first systematic characterization of intersectional demographic bias in safety refusals, proposing an evaluation framework grounded in targeted prompting and indirect adversarial attacks. Our methodology integrates refusal-rate analysis, response categorization, and statistical examination of refusal-length distributions. Empirical evaluation across mainstream LLMs reveals pervasive and statistically significant disparities in safety enforcement, exposing critical fairness gaps in current alignment and safety protocols. The study provides a reproducible methodological foundation and empirical evidence to guide the development of more robust, equitable AI safety strategies.

LLM safety guardrails exhibit selective refusal bias across demographicsModels inconsistently refuse harmful content for different demographic groupsSelective refusal enables indirect attacks targeting vulnerable demographic groups

Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models

Dec 09, 2024
NJ
Neel Jain
🏛️ University of Maryland | Capital One | New York University

Fine-grained, on-demand control of large language models’ refusal behavior remains challenging due to rigid, monolithic safety alignment. Method: We propose a plug-and-play refusal token mechanism that enables real-time, multi-dimensional, and fine-grained refusal rate control for a single pre-trained model—without any fine-tuning. Prior to generation, category-specific refusal tokens are injected; during inference, logit intervention and probability reweighting dynamically modulate their output probabilities, allowing precise refusal rate adjustment across sensitive query types (e.g., illegal, out-of-scope, or ambiguous queries). Results: Evaluated on multiple safety benchmarks, our method achieves ±1.2% refusal rate control error, supports instantaneous switching among user-defined refusal policies, and requires no additional training or model duplication. Its core contribution is the decoupling of refusal control from model parameters—establishing an efficient, lightweight paradigm for controllable safety alignment.

Calibrating refusal behavior in large language modelsControlling refusal rates without fine-tuning multiple modelsEnabling adjustable sensitivity to different query categories

Latest Papers

What's happening recently
View more

This work addresses inference-time refusal of large language models (LLMs) on politically sensitive topics. We propose Refusal Steering—a fine-grained, tuning-free control method that steers model activations to suppress undesired refusals while preserving safety and utility. Instead of hand-crafted rules, we employ an LLM-as-a-judge paradigm for automated refusal detection. A ridge-regularized steering vector is optimized in the activation space to precisely decouple refusal-oriented from compliant behavior. We empirically find that refusal signals are highly concentrated in deep Transformer layers and exhibit high-dimensional distributed patterns. On Qwen3-Next-80B, Refusal Steering eliminates excessive refusal on politically sensitive queries without degrading performance on JailbreakBench (safety) or standard general-purpose benchmarks—maintaining near-baseline accuracy. The method generalizes across model scales (4B and 80B variants) and enables bidirectional, real-time control over refusal behavior (i.e., on/off switching).

Control LLM refusal on sensitive topicsRemove political refusal while keeping safetyReplace pattern detection with LLM-as-judge

This study addresses the tendency of current language models to mechanically reject user requests even when confronted with unjust or absurd rules that warrant legitimate violation, revealing a critical deficiency in their capacity for moral judgment regarding rule legitimacy. The work introduces and systematically characterizes the phenomenon of “blind refusal,” constructing a synthetic dataset encompassing five categories of rule failure and nineteen authority types. Through automated quality control, human evaluation, and LLM-as-judge blind assessment using GPT-5.4, the authors conduct a two-dimensional behavioral analysis across 18 model configurations. Among 14,650 samples, models rejected 75.4% of exemption-eligible requests, with 57.5% of these cases involving correct identification of rule flaws yet still refusing assistance—demonstrating a pronounced disconnect between normative reasoning and behavioral decision-making.

blind refusaldefeated ruleslanguage models

Current large language models often erroneously reject safe inputs when defending against prompt injection attacks due to reliance on superficial heuristics rather than semantic intent. This work presents a systematic analysis of supervised fine-tuning–based defenses, identifying and quantifying three distinct shortcut behaviors: positional bias, token-trigger bias, and topic generalization bias. We construct targeted diagnostic datasets and conduct controlled experiments across multiple models and defense pipelines within a unified evaluation framework encompassing two base architectures. Our findings reveal that defended models exhibit up to 90% rejection rates on suffix-based inference tasks, with a single trigger token increasing false rejection rates by 50% and degrading test accuracy by as much as 40%. This study exposes fundamental limitations in current defense strategies and provides both diagnostic tools and empirical foundations for future research on robustness.

adversarial attacksfalse rejectionLLM security

This work addresses the challenge of coarse-grained safety alignment in current language models, which often leads to over-rejection of benign requests or under-rejection of harmful content. Focusing on Llama-3-8B, the study introduces category-specific refusal tokens and reveals, for the first time, that these tokens induce decoupled, category-aligned directions in the residual stream. Building on this insight, the authors propose a lightweight, training-free intervention: by constructing a unified, orthonormalized intervention vector via low-rank combinations in a whitened orthogonal basis, they enable precise, multi-category control over refusal behavior at inference time using probing techniques. This approach significantly reduces over-rejection on benign prompts while enhancing rejection of harmful queries, and demonstrates strong transferability across models sharing the same architecture.

category-specific refusallanguage model safetyover-refusal

Hot Scholars

LK

Lingkai Kong

Harvard University
Machine LearningData-Driven Decision MakingGenerative ModelsAI for Social Good
WZ

Wenwei Zhang

Shanghai AI Laboratory
Large Language ModelScalable OversightArtificial Intelligence
YG

Yuzhe Gu

Shanghai Jiao Tong University
Large Language ModelScalable OversightKnowledge and Reasoning
DL

Dahua Lin

The Chinese University of Hong Kong
computer visionmachine learningprobabilistic inferencebayesian nonparametrics
SB

Swarup Bhunia

University of Florida
IoT SecurityHardware SecurityEnergy-Efficient ElectronicsFood/Medicine Safety