🤖 AI Summary
This study addresses the security vulnerabilities arising from the cross-model transferability of adversarial attacks by proposing a lightweight defense method based on LoRA adapters. Leveraging the geometric properties of representation spaces, the approach employs contrastive learning with a frozen anchor model to compel the LoRA adapter to repel the internal representations of harmful prompts, thereby achieving effective defense without relying on computationally expensive adversarial training. Experimental results demonstrate that the proposed method reduces the cross-model attack success rate to below 1.1%, significantly outperforming existing defense baselines. Consequently, this work provides an efficient new paradigm for the safety alignment of large language models.
📝 Abstract
Adversarial attacks optimized on a single open-weight LLM can transfer to and jailbreak architecturally different models, allowing an attacker with white-box access to one model to compromise independently deployed systems. This creates a shared vulnerability across models, yet existing defenses are not designed for this cross-model threat. We find that cross-model transfer aligns with shared internal representation geometry, making it a natural defense target. AnchorRep targets this geometry directly with a lightweight LoRA adapter that pushes the defended model's internal representations of harmful prompts away from those of a frozen anchor model on the same prompts. Training uses a small set of harmful prompts and no adversarial examples. Across five models and four architectural families, AnchorRep reduces cross-model attack success rate to<=1.1% on 2,000 transferred attacks (0% on two), including the largest drop on Mistral (36% ->1.1%). Existing defenses can reduce transfer, but only at high cost either inducing up to 77% degenerate benign output or increasing over-refusal by up to 18%. Because such degenerate benign outputs are not captured by standard refusal-based metrics, we introduce the Benign Garble Rate to quantify them. Our results suggest that cross-model robustness can be achieved by shaping representation geometry, without requiring attack-specific training