Understanding and Enhancing Backdoor Persistency in LLM Agent Post-Training

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of backdoors in large language model agents to removal during post-training, which complicates the assessment of their persistence and associated supply-chain risks. To this end, it systematically analyzes the effects of supervised fine-tuning (SFT) and reinforcement learning (RL) on backdoor behavior and proposes PersistBD, a method that enhances backdoor survivability after benign training by optimizing gradient compatibility. Notably, the work reveals that the RL phase can preserve or even amplify residual backdoor behaviors. As the first effort to quantify and improve backdoor persistence throughout post-training, the proposed approach increases attack success rates from 20% to 74%–76% on Qwen2.5-Coder following SFT and SFT-RL pipelines, while maintaining utility on benign tasks.
📝 Abstract
Developers can build LLM agents by adapting third-party models through benign post-training. We study a supply-chain threat in which an attacker supplies a model with a backdoor: hidden behavior that produces malicious outputs when a particular input pattern appears. Focusing on software-engineering agents, we ask whether such backdoors survive the developer's supervised fine-tuning (SFT) and subsequent task-level reinforcement learning (RL). We observe that benign SFT substantially reduces attack success, but subsequent RL often preserves the residual behavior and sometimes even increases attack success. Our analysis of backdoor erosion during SFT identifies two factors that may favor survival: initial backdoor strength and gradient compatibility with benign training. These factors motivate PersistBD, which refines an already-backdoored model before release to improve its persistency through the benign post-training process. On Qwen2.5-Coder-7B, PersistBD raises attack success from 20% to 74% after SFT and from 20% to 76% after SFT-RL, while maintaining comparable benign task performance. Together, our results show that backdoors can remain active through benign post-training and that adversaries can deliberately increase their persistence. This highlights a supply-chain risk for AI developers and motivates stronger techniques for detecting and mitigating inherited backdoors when adapting third-party models into agents. Our code is available at https://github.com/uiuc-kang-lab/PersistBD.
Problem

Research questions and friction points this paper is trying to address.

backdoor persistency
supply-chain threat
LLM agents
post-training
adversarial robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Backdoor Persistency
Supply-chain Attack
Supervised Fine-tuning
Reinforcement Learning
PersistBD
🔎 Similar Papers
No similar papers found.