🤖 AI Summary
This study addresses the vulnerability of LoRA-fine-tuned large language models to backdoor attacks, noting that existing defenses rely on trigger priors, clean data, or retraining. To overcome these limitations, this work proposes an assumption-free, retraining-free purification method for LoRA parameters. By integrating data curation with feature approximation techniques, the approach constructs input-output orthogonal null-space projections to precisely eliminate backdoor directions embedded within LoRA weights. Experimental results demonstrate that the proposed method reduces the attack success rate from nearly 100% to below 10%, while fully preserving both the general capabilities of the base model and downstream task performance. To the best of our knowledge, this constitutes the first LoRA backdoor purification framework that operates without requiring any prior assumptions.
📝 Abstract
With the rapid adoption of large language models (LLMs) and parameter-efficient fine-tuning (PEFT) methods, the risk of backdoor attacks has become more severe. Existing backdoor purification methods typically rely on at least one of the strong assumptions, such as prior knowledge of triggers, access to clean references, or aggressive retraining, and they often lack comprehensive evaluations. These constraints substantially limit their practical applicability. To overcome these challenges, our work proposes purifying LoRA-tuned LLMs without these assumptions and even without post-hoc retraining of the suspect parameters. Our objective is to significantly reduce the attack success rates (ASR) while preserving both (i) the base model's general capabilities and (ii) the new downstream skills learned through the adapter. Through a series of ablation studies, we progressively scale our approach from a single layer in a text classification setting to a full-parameter LLM in the generative task. Through careful data curation and feature approximation, we extract high-fidelity backdoor directions and, for each layer or head, construct orthogonal null spaces in both the input and output channels, onto which the LoRA updates are projected. Empirically, our null-space projection method reduces the ASR from nearly 100% to less than 10%, while preserving the base model's benign performance and the adapter's learned abilities during downstream task adaptation.