Backdoor Mitigation in Decentralized LLM Fine-Tuning

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of decentralized large language model fine-tuning, where single-node poisoning can propagate backdoors to neighboring nodes via graph-based communication. To mitigate this threat, we propose Chorus, a defense mechanism that leverages local adapters as references to detect and reject compromised adapters prior to aggregation. Chorus operates through independent neighbor behavior probing and distributed collaborative voting, requiring neither shared validation data nor prior knowledge of attack triggers. Experimental results demonstrate that this approach reduces the average attack success rate on neighboring nodes to below 2.2%, closely approximating omniscient oracle performance while incurring negligible communication overhead. By effectively containing adversarial influence without centralized coordination, Chorus significantly enhances the security and robustness of decentralized federated learning frameworks against backdoor injection attacks.
📝 Abstract
Decentralized large language model (LLM) fine-tuning lets organizations collaboratively train a shared LLM on data they cannot pool, without a central coordinator. In every round, each node exchanges a trainable adapter with its neighbors over a communication graph, and then aggregates them. This setting, however, is vulnerable to propagated backdoors, which is a hidden behavior that lets a model perform normally on clean inputs but produce an attacker-chosen output whenever a secret trigger appears. We show that a single node poisoning its own model can backdoor adapters of nodes that have never seen a poisoned example, making them refuse prompts that contain a secret trigger. We present Chorus, a decentralized mechanism that lets each node detect and reject backdoored adapters from its neighbors before aggregation, without requiring shared validation data or any knowledge of the attacker's trigger or target. Chorus judges each adapter by its behavior, using the receiver's own adapter as a trusted reference. Crucially, no node in Chorus judges adapters alone: the receivers of each adapter update probe it independently, pool their findings in the neighborhood, and vote to make a decision. So a backdoor that slips past one receiver is still caught by the others. We evaluate the effectiveness of Chorus using two instruction-tuning datasets and LLM architectures, and against a state-of-the-art baseline. Chorus cuts the average attack success rate (ASR) of the attacker's neighbors from 48-63% to at most 2.2%, within 0.6 percentage points of an omniscient oracle that knows the exact malicious nodes. Even the worst-affected honest node never exceeds 10% ASR, the same bound as the oracle, against up to 78% without defense. This all comes at a negligible communication overhead.
Problem

Research questions and friction points this paper is trying to address.

Decentralized LLM fine-tuning
Backdoor attack
Poisoning propagation
Adapter aggregation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decentralized LLM Fine-Tuning
Backdoor Mitigation
Adapter Aggregation
Collaborative Voting
Behavior-based Detection