π€ AI Summary
This study addresses stealthy attacks in large language model fine-tuning services, where malicious data compromises safety alignment while preserving task performance. To counter this threat, we propose a lightweight defense mechanism integrating selective layer restoration with dynamic routing. Methodologically, by revealing the sign-based characteristics of inter-layer safety sensitivity, LoRA adapters are trained exclusively on extremal sensitive layers to restore safety capabilities. During inference, representation-level dynamic routing is introduced to precisely intercept malicious queries. Experimental results demonstrate that this approach substantially reduces harmful outputs while maintaining downstream task accuracy. Even under high poisoning rates, it sustains near-zero harm scores, achieving a favorable balance between safety and utility.
π Abstract
Fine-tuning-as-a-service enables users to adapt aligned large language models (LLMs) to specialized tasks, but malicious fine-tuning can erode refusal behavior while preserving task performance on legitimate inputs. We revisit recent layer-wise safety diagnostics and find that safety sensitivity is signed: scaling different layers can strengthen refusal, weaken it, or have little effect. Motivated by this observation, we propose SLDR, a post-fine-tuning defense based on Selective Layers Recovery and Dynamic Routing. SLDR trains a LoRA recovery adapter only on the layers with the maximum and minimum sensitivity scores in the signed spectrum, and uses representation-based dynamic routing inference to activate the adapter only for malicious queries. Across four model architectures, five downstream tasks, and four harmful benchmarks, SLDR substantially reduces harmful outputs while preserving downstream utility. On Llama3.1/SST2, SLDR reduces the average harmful score from 11.54 to 0.08 while maintaining downstream accuracy, and the harmful score remains near zero under poisoning ratios up to 0.9. The code is available at https://github.com/Stardust457/SLDR.