🤖 AI Summary
This study addresses the vulnerability of large language models to jailbreak attacks and the limitations of existing alignment methods, specifically superficial and excessive refusal. To overcome these issues, we propose the SSRFT framework, which introduces a novel safety role internalization paradigm grounded in psychometrics and role description, replacing conventional refusal-centric alignment strategies. By integrating supervised fine-tuning with synthetic data generation and multi-scenario expansion, we construct the SRQA dataset to shift model behavior from learning explicit refusal patterns toward deeply internalizing safety values. Experimental results demonstrate that our approach significantly enhances robustness against prefilling attacks and improves generalization to unseen jailbreak domains. Furthermore, it effectively reduces false refusal rates on benign queries while fully preserving general-purpose capabilities.
📝 Abstract
Large Language Models (LLMs) have achieved remarkable capabilities but remain vulnerable to jailbreak attacks that elicit harmful or unsafe outputs. Existing safety alignment approaches, including Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), often require substantial attack-specific supervision and computational resources, while remaining susceptible to shallow safety alignment and over-refusal. To address these challenges, we introduce SSRFT(Supervised Safe-Role Fine-Tuning), the first framework that reformulates safety alignment as the internalization of a predefined safe role. SSRFT constructs a Safe-Role Question-Answer (SRQA) dataset from psychometric questions, limited jailbreak prompts, and a safe-role description. Role-consistent responses are synthesized, validated, and expanded into diverse scenarios, enabling models to internalize safety-oriented values and principles rather than explicit refusal patterns. Experiments across multiple Base and Instruct models show that SSRFT achieves more robust and generalizable safety alignment than standard SFT. SSRFT shows substantially greater robustness to prefilling attacks and better generalization to unseen jailbreak domains, while reducing over-refusal on benign queries and preserving the model's general capabilities. These results establish safe-role internalization as an effective alternative to refusal-centric safety alignment. Warning: This paper contains examples of harmful and toxic language.