Theory of Minimal Weight Perturbations in Deep Networks and its Applications for Low-Rank Activated Backdoor Attacks

📅 2026-01-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates how to quantify the minimal weight perturbation required to induce a specified output change in deep neural networks, with applications to analyzing and defending against backdoor attacks that exploit low-rank compressed activations. We derive, for the first time, closed-form expressions for the minimum-norm weight perturbations in both single-layer and multi-layer networks, revealing how backpropagation margins govern layer-wise sensitivity and establishing connections to Lipschitz-based robustness bounds. Building on these insights, we propose provable compression thresholds that delineate the feasibility boundary of backdoor attacks. Theoretical analysis and experiments demonstrate that low-rank compression can reliably trigger latent backdoors while preserving model accuracy, and simultaneously provide certifiable guarantees on the minimality of parameter updates needed for such attacks.

Technology Category

Computer Vision: Adversarial Attacks & RobustnessMachine Learning: Adversarial Learning & RobustnessSearch and Optimization: Learning to Search

Application Category

Graph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphsWeb Mining and Content Analysis: Robustness and generalizability of Web mining methodsResponsible Web: Measurement, analysis, and circumvention of Web censorship
📝 Abstract
The minimal norm weight perturbations of DNNs required to achieve a specified change in output are derived and the factors determining its size are discussed. These single-layer exact formulae are contrasted with more generic multi-layer Lipschitz constant based robustness guarantees; both are observed to be of the same order which indicates similar efficacy in their guarantees. These results are applied to precision-modification-activated backdoor attacks, establishing provable compression thresholds below which such attacks cannot succeed, and show empirically that low-rank compression can reliably activate latent backdoors while preserving full-precision accuracy. These expressions reveal how back-propagated margins govern layer-wise sensitivity and provide certifiable guarantees on the smallest parameter updates consistent with a desired output shift.
Problem

Research questions and friction points this paper is trying to address.

minimal weight perturbations
deep neural networks
backdoor attacks
low-rank compression
output sensitivity
Innovation

Methods, ideas, or system contributions that make the work stand out.

minimal weight perturbation
low-rank compression
backdoor activation
Lipschitz robustness
certifiable guarantees
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.