🤖 AI Summary
This study addresses the limited stealthiness of existing training-time attacks, which predominantly rely on tampering with labels or model weights. To overcome this limitation, we propose a parameter-free availability attack paradigm that exploits the gradient blocking mechanism induced by ReLU neuron death. Specifically, this work introduces a novel strategy combining dynamic data ordering with gradient inversion-based poisoning. By employing greedy sorting to select samples and synthesizing adversarial data to deactivate target neurons, the proposed method effectively disrupts training without modifying parameters. Furthermore, soft knock-out weight reordering is incorporated to amplify the destructive impact. Experimental evaluations on MNIST demonstrate that merely a small number of reordered or poisoned samples can significantly degrade test accuracy from 96% to 86%–88%, thereby validating both the effectiveness and the severe threat posed by this highly covert attack framework.
📝 Abstract
Rectified linear unit (ReLU) networks can suffer from dying neurons, where units with persistently negative pre-activations produce zero outputs, blocking gradients through their activations. To exploit this failure mode, we present three training-time availability attacks based on data ordering and poisoning. We begin with the basic dynamic data-ordering attack (DOA), which greedily constructs a training prefix by selecting the next example that minimizes the target layer's post-update weight sum, aiming to push ReLU units toward negative pre-activations without modifying training samples or labels. We then develop two poisoning attacks, IG-DOA and IG-SKA, which use gradient inversion to synthesize class-conditioned samples by matching reference gradients in adverse model states constructed through data ordering or soft knockout, respectively. Soft knockout rearranges weights across adjacent layers to concentrate negative contributions. On a fully connected ReLU network trained on MNIST, ordering 100 of 60,000 training examples reduces test accuracy from 96% to 95% after only five epochs. Adding 200 poisoned samples from a single class reduces test accuracy to approximately 86-88% after five epochs in most evaluated conditions, compared with approximately 96% under clean training. These results demonstrate that ReLU-targeted data ordering and poisoning can impair learning without directly modifying the victim model's parameters.