Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks

πŸ“… 2026-08-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the vulnerability of large foundation models to white-box attacks targeting internal safety routing in open-world settings, where existing defenses are brittle due to their reliance on static safety mechanisms. To overcome this limitation, we propose the Dynamic Routing Adaptive Alignment (DRAA) framework, which introduces a novel dynamic compensation routing mechanism. DRAA activates contrastive localization to identify and mask safety-critical routing pathways, thereby inducing causal failure modes. These failures are leveraged to construct preference pairs, which in turn drive the dynamic reconfiguration of rejection pathways. By restructuring the model’s dependency on fixed safety routes, DRAA significantly enhances robustness against routing-level white-box attacks while preserving strong general-purpose capabilities.
πŸ“ Abstract
With the widespread deployment of large foundation models (LFMs) in open environments, safety threats are shifting from black-box jailbreaks toward white-box attacks that directly identify and disrupt internal safety neurons or routes. However, existing safety defenses often rely on static safety units or fixed refusal pathways, leaving models highly vulnerable to targeted route-level white-box attacks. For that, we propose dynamic routing adaptive alignment (DRAA), a framework that introduces dynamic compensatory routes to preserve robust refusal behavior when the safety route is compromised. Specifically, we first identify and localize the model's safety route by contrasting internal activations between safe and unsafe calibration samples. DRAA then masks this safety route to induce causal failure cases and selectively mines the resulting defense failures, thereby constructing failure-aware preference pairs. Extensive experiments demonstrate that DRAA effectively restructures the underlying pathway dependence of model safety, substantially improving robustness against route-level white-box attacks, while preserving general utility.
Problem

Research questions and friction points this paper is trying to address.

white-box attacks
safety alignment
dynamic routing
foundation models
route-level vulnerability
Innovation

Methods, ideas, or system contributions that make the work stand out.

dynamic routing
adaptive alignment
white-box attacks
safety neurons
failure-aware preference
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
S
Shangze Li
Nanjing University of Science and Technology
C
Chuancheng Shi
The University of Sydney
S
Simiao Xie
The University of Sydney
L
Lingzhi He
University of New South Wales
C
Cheng Ji
Nanjing University of Science and Technology
Z
Zifeng Cheng
Nanjing University
Fei Shen
Fei Shen
National University of Singapore
Controllable GenerationMultimodal Safety
C
Chao Wu
Nanjing University of Science and Technology
Tat-Seng Chua
Tat-Seng Chua
National University of Singapore
Multimedia Information RetrievalLive Social Media Analysis