🤖 AI Summary
This study addresses the computational intractability of directly evaluating probabilistic robustness and the inefficiency of existing adversarial example generation and defense mechanisms. Motivated by the intuition of distributional overlap, the authors derive a lower bound on the Kullback-Leibler divergence as a tractable surrogate objective and propose a novel probabilistic adversarial training algorithm. Furthermore, this work provides a probabilistic interpretation of conventional adversarial training, formally proving its equivalence to maximizing a lower bound on robustness. The proposed framework introduces a new theoretical perspective on adversarial robustness and significantly enhances the probabilistic robustness of deep learning models. Notably, the introduced scaling factor can be seamlessly integrated into existing non-probabilistic adversarial training methods to improve their performance without architectural modifications.
📝 Abstract
Building on a probabilistic perspective in which adversarial examples arise from the overlap between a distance-based distribution $p_{\mathrm{dis}}$ and a victim-classifier-induced distribution $p_{\mathrm{vic}}$, we start from a simple intuition: adversarial examples become harder to generate when these two distributions are pushed apart, as their overlap becomes smaller, thereby increasing robustness. This intuition naturally motivates a KL-based robustness objective. We then prove that $\mathrm{KL}(p_{\mathrm{dis}}\|p_{\mathrm{vic}})-\log Z_{\mathrm{vic}}$ is a lower bound on probabilistic robustness (PR), where $Z_{\mathrm{vic}}$ denotes the normalizing constant of $p_{\mathrm{vic}}$. Since PR is generally intractable to compute directly, maximizing this KL-based lower bound provides a tractable surrogate objective for improving PR. We further show that this objective recovers a scaled form of adversarial training, offering a probabilistic interpretation of adversarial training and a principled route to robustness improvement. We call the resulting method probabilistic adversarial training. Experiments show that it consistently improves PR, and ablation studies demonstrate that the induced scaling factor can even enhance the PR of non-probabilistic adversarial training methods.