🤖 AI Summary
Adversarial examples in black-box attacks often exhibit weak transferability across models. Method: This paper proposes BayAtk, a Bayesian inference-based transferable adversarial attack method. It introduces Bayesian modeling to adversarial transferability analysis for the first time, uncovering the inherent uncertainty in transferability. BayAtk designs two transferability-enhancing priors and incorporates an instance-adaptive dynamic weighting mechanism to jointly optimize perturbation direction and magnitude. Contribution/Results: Evaluated on undefended and state-of-the-art defended black-box models, BayAtk significantly outperforms existing SOTA methods, achieving an average 12.7% improvement in transfer success rate. It establishes a more robust and interpretable attack paradigm for AI security evaluation.
📝 Abstract
The vulnerability of deep neural networks (DNNs) to black-box adversarial attacks is one of the most heated topics in trustworthy AI. In such attacks, the attackers operate without any insider knowledge of the model, making the cross-model transferability of adversarial examples critical. Despite the potential for adversarial examples to be effective across various models, it has been observed that adversarial examples that are specifically crafted for a specific model often exhibit poor transferability. In this paper, we explore the transferability of adversarial examples via the lens of Bayesian approach. Specifically, we leverage Bayesian approach to probe the transferability and then study what constitutes a transferability-promoting prior. Following this, we design two concrete transferability-promoting priors, along with an adaptive dynamic weighting strategy for instances sampled from these priors. Employing these techniques, we present BayAtk. Extensive experiments illustrate the significant effectiveness of BayAtk in crafting more transferable adversarial examples against both undefended and defended black-box models compared to existing state-of-the-art attacks.