Robustness Feature Adapter for Efficient Adversarial Training

πŸ“… 2025-08-25
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
To address the dual challenges of high computational cost and robust overfitting in adversarial training (AT) for large backbone models, this paper proposes Feature-space Adapter-based Adversarial Training (FA-AT). FA-AT embeds lightweight adapter modules into intermediate feature layers, shifting adversarial perturbation injection and robust optimization from the input space to the feature spaceβ€”thereby eliminating repeated, expensive gradient computations on raw inputs. Leveraging the low-rank structure of adapters, FA-AT implicitly regularizes the robust learning process, mitigating overfitting. Integrated with PGD, FA-AT achieves 35–52% training speedup on ResNet and ViT backbones, while improving robust accuracy by 2.1–4.8 percentage points on average. Moreover, it significantly enhances generalization to unseen attacks.

Technology Category

Computer Vision: Adversarial Attacks & RobustnessMachine Learning: Adversarial Learning & RobustnessMultiagent Systems: Adversarial Agents

Application Category

User Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systemsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingWeb Mining and Content Analysis: Large pretrained models with web data
πŸ“ Abstract
Adversarial training (AT) with projected gradient descent is the most popular method to improve model robustness under adversarial attacks. However, computational overheads become prohibitively large when AT is applied to large backbone models. AT is also known to have the issue of robust overfitting. This paper contributes to solving both problems simultaneously towards building more trustworthy foundation models. In particular, we propose a new adapter-based approach for efficient AT directly in the feature space. We show that the proposed adapter-based approach can improve the inner-loop convergence quality by eliminating robust overfitting. As a result, it significantly increases computational efficiency and improves model accuracy by generalizing adversarial robustness to unseen attacks. We demonstrate the effectiveness of the new adapter-based approach in different backbone architectures and in AT at scale.
Problem

Research questions and friction points this paper is trying to address.

Reducing computational overhead in adversarial training for large models
Eliminating robust overfitting to improve convergence quality
Generalizing adversarial robustness to unseen attacks efficiently
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adapter-based adversarial training in feature space
Eliminates robust overfitting through improved convergence
Enhances computational efficiency and attack generalization
πŸ”Ž Similar Papers
No similar papers found.
Q
Quanwei Wu
Dongguan University of Technology
J
Jun Guo
Dongguan University of Technology
W
Wei Wang
The Hong Kong University of Science and Technology (Guangzhou)
Y
Yi Wang
Dongguan University of Technology