Refusal Localizes, the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of large language models to safety alignment degradation via fine-tuning on minimal harmful samples, noting that existing localization-based defenses are susceptible to adaptive attacks. We systematically evaluate the robustness of strategies such as layer freezing and spectral repair under dynamic adversarial conditions. By integrating linear separability analysis, hidden state patching, and singular value decomposition, this work reveals a "damage location migration" phenomenon during fine-tuning. We demonstrate that adversaries can shift target regions to circumvent cross-model repairs, thereby proposing five criteria for defense evaluation. Our results indicate that although localized repair struggles against adaptive attacks, layer freezing retains practical utility in mitigating non-malicious data leakage scenarios.
📝 Abstract
Fine-tuning adapts aligned large language models (LLMs) to downstream tasks, but a few dozen harmful examples can remove their refusal of harmful requests. Prior work localizes safety-related behavior to specific layers, directions, and tokens, suggesting targets for protection. We test whether successful localization and recovery support defenses that survive changes in the attack. Across six checkpoints from four model families, harmful and benign prompts remain linearly separable after attack, and patching full clean hidden states into the compromised model restores refusal at a reproducible transition depth. Building on a prior layer-freezing defense, we freeze every layer up to this depth and repeat the attack. At a hundred harmful examples, refusal remains near zero on all six checkpoints, with recovery transitions above the frozen boundary. In a second study, removing the update's top two singular directions restores refusal after short attention-only fine-tunes on four checkpoints. On Llama-3.1-8B, ordinary training changes weaken this repair and an attacker who spreads the update defeats it. A spectral detector calibrated on benign Llama fine-tunes misses most repair failures on that checkpoint. Localized freezing can nevertheless help preserve refusal when a few harmful examples enter training data unintentionally. These results show that an attacker can bypass a region identified by recovery and defeat a repair that works across multiple checkpoints, motivating five checks for defenses against adaptive fine-tuning. Code is available at https://github.com/js-lee-AI/refusal-relocates.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Safety Alignment
Fine-Tuning
Adversarial Attack
Refusal Mechanism
Innovation

Methods, ideas, or system contributions that make the work stand out.

safety alignment
few-sample fine-tuning
refusal localization
singular value decomposition
adaptive attack