Safer Content or Firmer Refusals? A Hybrid Perturbation Defense for Alignment under Harmful Fine-tuning

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the alignment breakdown of large language models caused by harmful fine-tuning by proposing VaccineBooster, a hybrid defense framework. To the best of our knowledge, this method is the first to jointly integrate embedding-layer perturbation with weight-layer gradient attenuation, achieving single-step collaborative training through the simulation of adversarial updates to effectively mitigate data poisoning attacks. Furthermore, this work reveals an intrinsic trade-off mechanism between content safety and refusal behavior. Experimental evaluations demonstrate that the proposed framework achieves the lowest violation score of 0.315 on Llama-2, while also providing practical guidelines for configuring safety strategies in aligned language models.
📝 Abstract
Fine-tuning-as-a-service lets users adapt a safety-aligned language model to their own data, but it also creates a harmful fine-tuning attack surface: a small amount of harmful data mixed into an otherwise benign fine-tuning set can degrade the model's alignment. Two recent alignment-stage defenses address this problem at different levels of the model. Vaccine improves the robustness of hidden embeddings to the representation shifts induced by harmful fine-tuning, whereas Booster simulates harmful weight updates and attenuates their effect during alignment. We investigate whether these mechanisms are complementary and propose VaccineBooster, a single alignment procedure that combines embedding perturbation and weight-level gradient attenuation within each training step. On Llama-2-7B aligned with BeaverTails and then attacked through poisoned fine-tuning, VaccineBooster achieves the lowest OpenAI moderation score among the compared defenses, 0.315, while a Booster-Only variant retains the highest post-attack refusal rate, 50%. Together with ablations over the embedding-perturbation and gradient-attenuation strengths, these results indicate a trade-off: embedding perturbation primarily reduces flagged harmful content, whereas gradient attenuation primarily preserves explicit refusal behavior. Because our evaluation uses ten prompts and a single unseeded run per configuration, we report this trade-off as an observed pattern rather than a statistically resolved effect. These results provide practical guidance for prioritizing content safety or refusal retention when aligned models are exposed to untrusted fine-tuning.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Harmful Fine-tuning Defense
Hybrid Perturbation
Embedding Perturbation
Gradient Attenuation
Alignment Robustness
M
Muhammad Zeeshan Akram
University of Louisville
M
Mufid Kamel Marican
University of Louisville
A
Anvesh Reddy Yenugu
University of Louisville
A
Ali Zain Kaimkhani
University of Louisville
Minghong Fang
Minghong Fang
University of Louisville
SecurityPrivacyAI SafetyMachine Learning