π€ AI Summary
To address the dual challenges of high computational cost and robust overfitting in adversarial training (AT) for large backbone models, this paper proposes Feature-space Adapter-based Adversarial Training (FA-AT). FA-AT embeds lightweight adapter modules into intermediate feature layers, shifting adversarial perturbation injection and robust optimization from the input space to the feature spaceβthereby eliminating repeated, expensive gradient computations on raw inputs. Leveraging the low-rank structure of adapters, FA-AT implicitly regularizes the robust learning process, mitigating overfitting. Integrated with PGD, FA-AT achieves 35β52% training speedup on ResNet and ViT backbones, while improving robust accuracy by 2.1β4.8 percentage points on average. Moreover, it significantly enhances generalization to unseen attacks.
π Abstract
Adversarial training (AT) with projected gradient descent is the most popular method to improve model robustness under adversarial attacks. However, computational overheads become prohibitively large when AT is applied to large backbone models. AT is also known to have the issue of robust overfitting. This paper contributes to solving both problems simultaneously towards building more trustworthy foundation models. In particular, we propose a new adapter-based approach for efficient AT directly in the feature space. We show that the proposed adapter-based approach can improve the inner-loop convergence quality by eliminating robust overfitting. As a result, it significantly increases computational efficiency and improves model accuracy by generalizing adversarial robustness to unseen attacks. We demonstrate the effectiveness of the new adapter-based approach in different backbone architectures and in AT at scale.