🤖 AI Summary
This study addresses the limited cross-model reusability of existing adversarial example detectors caused by their dependence on specific classification backbones. We propose a graph-based universal detection framework that leverages graph neural networks to transform intermediate-layer features into structured representations. By introducing a cross-backbone representation alignment technique, our approach overcomes the bottleneck of ineffective transfer of detection knowledge across heterogeneous architectures, enabling direct detector reuse without training from scratch. Extensive evaluations across multiple datasets and adaptive attack scenarios demonstrate that the proposed method achieves significantly superior aggregated ROC-AUC compared to both from-scratch training and existing transfer baselines. This work establishes a novel paradigm for constructing plug-and-play, highly generalizable adversarial defense systems.
📝 Abstract
Adversarial example detectors are often tied to the classifier backbone they were trained on, limiting reuse when the protected model is replaced or upgraded. Directly transferring such detectors across backbones is challenging because different networks generally produce incompatible internal representations. We propose GraphRectify, a graph-based framework for transferring adversarial image detectors across classifier backbones. GraphRectify learns a structured representation of intermediate classifier features and adapts representations from a new backbone to the detector learned on the original model, enabling detector reuse. We evaluate GraphRectify across multiple datasets, backbone architectures, and adversarial attacks, including detector-aware adaptive attacks that jointly target the classifier and detector. Across the complete evaluation matrix, GraphRectify achieves higher aggregate ROC-AUC than training a detector from scratch on the new backbone and the evaluated transfer ablations. The gains are particularly strong for transfers between different backbone families and when sufficient data are available. In contrast, training from scratch remains competitive in the most data-limited settings. These results show that adversarial detection knowledge can transfer effectively across heterogeneous classifier architectures rather than being relearned whenever the protected backbone changes.