Target-free Latent Safety Alignment

πŸ“… 2026-10-03
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation of existing adversarial training methods, which rely on fixed harmful targets and thus exhibit homogeneous behavioral patterns that struggle to defend against diverse jailbreak attacks. To overcome this, we propose a targetless adversarial training framework that eliminates predefined objectives. By leveraging latent-space behavioral shift amplification and semantic entropy measurement, our approach induces and generates highly diverse adversarial examples in an unsupervised manner. This method effectively overcomes the narrow behavioral coverage inherent in conventional training paradigms. Consequently, it significantly enhances the safety alignment robustness and overall defensive capabilities of large language models against various unseen jailbreak attacks.
πŸ“ Abstract
Large language models (LLMs) remain highly vulnerable to jailbreak attacks that induce harmful behaviors and circumvent safety alignment. To defend against such attacks, adversarial training paradigms have been proposed to first simulate failure modes and then train the model to correct them, yielding promising improvements in safety alignment. However, these methods typically construct adversarial samples either by encouraging fixed harmful target completions or by performing targeted activation ablation derived from fixed benign--harmful data pairs. As a result, the generated adversarial samples tend to induce homogeneous harmful behaviors that poorly reflect the diversity of behaviors elicited by real-world jailbreak attacks. This behavior-level narrowness fundamentally limits their robustness. To address this issue, we propose a target-free adversarial training framework that generates adversarial samples in an unsupervised manner. By amplifying and diversifying behavior-level shifts in the model's latent space, our approach produces semantically diverse adversarial samples that induce a wide range of harmful behaviors. This expanded behavioral coverage exposes more diverse failure modes and thereby improves safety alignment. To quantify this effect, we use semantic entropy as an output-level measure of adversarial behavioral diversity. Empirically, our method elicits diverse harmful behaviors in the target model, substantially mitigating behavioral narrowness and improving robustness to jailbreak attacks.
Problem

Research questions and friction points this paper is trying to address.

jailbreak attacks
safety alignment
adversarial training
behavioral narrowness
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Target-free Adversarial Training
Latent Safety Alignment
Jailbreak Robustness
Semantic Entropy
Unsupervised Adversarial Generation
πŸ”Ž Similar Papers
L
Luoyu Chen
University of Technology Sydney, Sydney, NSW, Australia
W
Weiqi Wang
Xi’an Jiaotong University, Xi’an, Shaanxi, China
Chenhan Zhang
Chenhan Zhang
PhD
deep Learningprivacy-preserving
Z
Zhiyi Tian
Southeast University, Nanjing, Jiangsu, China
Y
Yuxian Huang
Nanjing University of Posts and Telecommunications, Nanjing, Jiangsu, China
S
Shui Yu
University of Technology Sydney, Sydney, NSW, Australia