Explaining the Saliency Map Sparsity of Adversarially-Trained Neural Networks

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates why saliency maps of adversarially trained neural networks exhibit sparsity. Focusing on two-layer ReLU networks, we conduct an asymptotic analysis within the empirical risk minimization framework under $L_\infty$-bounded adversarial attacks. We provide the first theoretical proof that adversarial training is equivalent to weight decay combined with total variation regularization, and converges to a Bayes classifier minimizing the gradient norm. This reveals the intrinsic mechanism by which gradient norm minimization induces axis-aligned sparsity. Furthermore, experiments utilizing the gradient $L_1$ norm and thresholded sparsity metrics confirm that saliency maps from adversarially trained models are substantially sparser than those from naturally trained counterparts. These findings offer rigorous theoretical support for understanding adversarial robustness.
📝 Abstract
Understanding why deep neural networks make a given prediction is of great importance for their safe deployment. In computer vision, saliency maps, which highlight the image region most influential for a prediction, remain a widely-used form of explanation. An empirical observation is the apparent sparsity of gradient saliency maps of adversarially-trained neural networks. In this paper, we propose a theoretical explanation of this phenomenon for two-layer ReLU networks. We build on the established equivalence of adversarial training to the minimization of the empirical risk with weight-decay penalization and an added adversarial total variation term -- valid for certain loss functions. As the number of data points and neurons grows and the regularization parameters are sent to zero at appropriate rates, we prove that minimizers converge to a Bayes classifier with minimal gradient and Barron norm. Sparsity appears since for adversarial training with $\ell_\infty$-attacks the gradient norm is anisotropic and favors axis-aligned / sparse gradients. We illustrate our theoretical findings experimentally by evaluating the gradient $\ell_1$-norm and thresholded sparsity of naturally versus adversarially trained models.
Problem

Research questions and friction points this paper is trying to address.

adversarial training
saliency map sparsity
deep neural networks
gradient norm
explainability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adversarial Training
Saliency Map Sparsity
Gradient Norm Anisotropy
Barron Norm
Bayes Classifier
🔎 Similar Papers
No similar papers found.