Why Backdooring Neural Networks is so Easy?

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unresolved theoretical question of why existing linear auditing methods systematically underestimate the backdoor vulnerability of feature-learning models. By deriving closed-form analytical solutions for quadratic neurons trained on Gaussian mixture data, this work contrasts lazy and feature learning dynamics to establish, for the first time, that the required attack budget scales with the poisoning ratio according to a π^{-1/4} law. Theoretically, it demonstrates that while feature learning enhances predictive performance, it simultaneously exacerbates backdoor vulnerabilities by substantially reducing the necessary trigger strength. These findings reveal how nonlinear mechanisms amplify security risks, providing a principled explanation for large-scale empirical observations and confirming that conventional linear heuristic auditing fails under prevailing feature-learning paradigms.
📝 Abstract
Securing modern AI systems against backdoor attacks remains an open challenge and requires fundamentally principled estimates of the adversary's budget -- the poison fraction $π$ and trigger strength $α$ needed to construct successful yet stealthy attacks. Motivated by recent empirical evidence that poisoning large language models can require a nearly constant number of malicious samples even as clean datasets grow, we derive an exact closed-form analysis of a quadratic neuron trained on a poisoned Gaussian mixture. We show, perhaps counterintuitively, that the same feature-learning dynamics that make neural networks powerful can also make them more vulnerable to backdoors. Specifically, with clean accuracy preserved to first order, $O(π)$, we demonstrate that lazy learning imposes the inverse-square-root scaling $α\propto π^{-1/2}$ for a successful attack, while feature learning induces a quadratic detector whose loss margin scales as $O(α^4)$, improving the attack budget to $α\propto π^{-1/4}$. Consequently, nonlinear feature learning substantially reduces the trigger strength required at small poison fractions, thereby in a sense making feature learners more backdoor vulnerable. These results provide a theoretical mechanism consistent with large-scale empirical observations and demonstrate that security audits based on linear heuristics can systematically underestimate backdoor vulnerability in the widely adopted feature-learning regimes.
Problem

Research questions and friction points this paper is trying to address.

backdoor attacks
poison fraction
trigger strength
feature learning
neural network vulnerability
Innovation

Methods, ideas, or system contributions that make the work stand out.

backdoor attacks
feature learning
scaling law
quadratic neuron
Gaussian mixture