The Model Plants the Trigger: Answer-Side Backdoor Attacks in Multi-Turn Large Language Models

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses a critical blind spot in existing LLM backdoor defenses, which focus solely on input triggers while overlooking risks from model-generated content. We propose an answer-side backdoor attack in multi-turn dialogues, introducing the novel paradigm of "model-planted triggers." Specifically, a benign initial turn induces the model to output specific tokens that enter the conversation history; these self-generated tokens are subsequently exploited to bypass safety alignment when harmful queries follow. The stealthy implantation is achieved through data poisoning, representation-level analysis, and adversarial prompt engineering. Experiments across four LLMs demonstrate that a mere 5% poisoning rate yields near-perfect attack success while preserving general utility. These findings expose fundamental vulnerabilities in current input-sanitization defense frameworks.
📝 Abstract
Safety alignment in Large Language Models (LLMs) remains vulnerable to backdoor attacks. Existing LLM backdoors are almost all input-centric: activation depends on explicit trigger patterns in the user input, so modern guardrails are built to sanitize the input space. We challenge this assumption with a novel answer-side backdoor for multi-turn dialogue. Instead of inserting the trigger into the input, the adversary uses a benign first-turn prompt to naturally induce the model to generate a specific, seemingly innocuous word. Once merged into the dialogue history, this self-generated word becomes the trigger. When a later harmful query arrives, the model detects its own trigger and bypasses its safety refusal, while the user input stays perfectly clean. Across four LLMs, our attack reaches near-perfect Attack Success Rates, approaching 100\% at only a 5\% poisoning rate, while preserving general utility and clean-input safety, and it evades mainstream input-centric defenses. Representation-level analysis shows that the self-generated trigger consistently suppresses the model's refusal signal, exposing a critical blind spot in current LLM defenses.
Problem

Research questions and friction points this paper is trying to address.

Backdoor Attacks
Large Language Models
Safety Alignment
Multi-Turn Dialogue
Answer-Side Trigger
Innovation

Methods, ideas, or system contributions that make the work stand out.

Answer-side backdoor
Multi-turn dialogue
Self-generated trigger
Safety alignment
Large Language Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yibo Zhang
Queen Mary University of London, UK
T
Tianrong Guan
Squirrel AI Learning, USA
Liang Lin
Liang Lin
Fellow of IEEE/IAPR, Professor of Computer Science, Sun Yat-sen University
Embodied AICausal Inference and LearningMultimodal Data Analysis
P
Puze Wang
Queen Mary University of London, UK
J
Jin Wang
Squirrel AI Learning, USA
Q
Qingsong Wen
Squirrel AI Learning, USA