🤖 AI Summary
This study addresses the challenge in large language model backdoor defense where backdoor removal often induces output distribution shifts and performance degradation. We propose NEEDLE, a method that requires neither a clean reference model nor the original poisoned data. By estimating the backdoor direction and rejection subspace via activation vectors, NEEDLE performs training-free model editing through weight orthogonalization to precisely suppress backdoors while preserving benign representations. Experiments demonstrate that NEEDLE achieves the lowest average attack success rate—reaching 0% for code injection—and minimal KL divergence across various models and attack settings. Furthermore, it effectively prevents performance degradation on benign prompts, successfully balancing model utility and security.
📝 Abstract
Backdoor attacks can be implanted in Large Language Models (LLMs) during training, causing unwanted behaviour when a trigger appears in the input. Existing backdoor defences for LLMs attempt to remove the backdoor but inadvertently shift the model's output distribution to benign prompts, which can result in degraded model performance and safety. We propose NEEDLE, a training-free method for targeted backdoor removal. Once a trigger has been identified, our method estimates a backdoor direction and a refusal subspace through activation vectors, then applies sequential weight orthogonalisation to suppress the backdoor while preventing changes in refusal-related representations. NEEDLE requires neither a clean reference model nor the original poisoned training data. Evaluation is conducted across multiple model families and attack types. NEEDLE achieves the lowest mean Attack Success Rate (ASR) among the evaluated defences, including 0% on challenging code injection attacks, while resulting in the lowest KL divergence and minimal changes in capability and safety.