π€ AI Summary
This study addresses the limitation of existing adversarial training methods, which rely on fixed harmful targets and thus exhibit homogeneous behavioral patterns that struggle to defend against diverse jailbreak attacks. To overcome this, we propose a targetless adversarial training framework that eliminates predefined objectives. By leveraging latent-space behavioral shift amplification and semantic entropy measurement, our approach induces and generates highly diverse adversarial examples in an unsupervised manner. This method effectively overcomes the narrow behavioral coverage inherent in conventional training paradigms. Consequently, it significantly enhances the safety alignment robustness and overall defensive capabilities of large language models against various unseen jailbreak attacks.
π Abstract
Large language models (LLMs) remain highly vulnerable to jailbreak attacks that induce harmful behaviors and circumvent safety alignment. To defend against such attacks, adversarial training paradigms have been proposed to first simulate failure modes and then train the model to correct them, yielding promising improvements in safety alignment. However, these methods typically construct adversarial samples either by encouraging fixed harmful target completions or by performing targeted activation ablation derived from fixed benign--harmful data pairs. As a result, the generated adversarial samples tend to induce homogeneous harmful behaviors that poorly reflect the diversity of behaviors elicited by real-world jailbreak attacks. This behavior-level narrowness fundamentally limits their robustness. To address this issue, we propose a target-free adversarial training framework that generates adversarial samples in an unsupervised manner. By amplifying and diversifying behavior-level shifts in the model's latent space, our approach produces semantically diverse adversarial samples that induce a wide range of harmful behaviors. This expanded behavioral coverage exposes more diverse failure modes and thereby improves safety alignment. To quantify this effect, we use semantic entropy as an output-level measure of adversarial behavioral diversity. Empirically, our method elicits diverse harmful behaviors in the target model, substantially mitigating behavioral narrowness and improving robustness to jailbreak attacks.