Leveraging Generalizability of Image-to-Image Translation for Enhanced Adversarial Defense

📅 2025-04-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the vulnerability of image classification models to diverse unknown adversarial attacks and the poor generalizability and high computational overhead of existing defenses, this paper proposes a lightweight preprocessing defense framework based on image-to-image translation. Our method trains only a single residual-enhanced translation model, enabling robust adversarial purification—without fine-tuning—across attack types (e.g., FGSM, PGD, CW) and target models (e.g., ResNet, VGG, ViT). The key innovation lies in incorporating a residual architecture into the translation model to enhance cross-domain generalization, eliminating the need for attack- or model-specific customization. On a multi-attack–multi-model benchmark, our approach restores classification accuracy from near 0% to an average of 72%, while incurring significantly lower inference latency and memory footprint compared to state-of-the-art defense methods.

Technology Category

Computer Vision: Adversarial Attacks & RobustnessMachine Learning: Adversarial Learning & RobustnessNatural Language Processing: Safety and Robustness

Application Category

User Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systemsWeb Mining and Content Analysis: Large pretrained models with web dataSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web search
📝 Abstract
In the rapidly evolving field of artificial intelligence, machine learning emerges as a key technology characterized by its vast potential and inherent risks. The stability and reliability of these models are important, as they are frequent targets of security threats. Adversarial attacks, first rigorously defined by Ian Goodfellow et al. in 2013, highlight a critical vulnerability: they can trick machine learning models into making incorrect predictions by applying nearly invisible perturbations to images. Although many studies have focused on constructing sophisticated defensive mechanisms to mitigate such attacks, they often overlook the substantial time and computational costs of training and maintaining these models. Ideally, a defense method should be able to generalize across various, even unseen, adversarial attacks with minimal overhead. Building on our previous work on image-to-image translation-based defenses, this study introduces an improved model that incorporates residual blocks to enhance generalizability. The proposed method requires training only a single model, effectively defends against diverse attack types, and is well-transferable between different target models. Experiments show that our model can restore the classification accuracy from near zero to an average of 72% while maintaining competitive performance compared to state-of-the-art methods.
Problem

Research questions and friction points this paper is trying to address.

Enhancing adversarial defense with image-to-image translation
Improving model generalizability against diverse adversarial attacks
Reducing training overhead while maintaining defense effectiveness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses image-to-image translation for defense
Incorporates residual blocks for generalizability
Trains single model for diverse attacks
H
Haibo Zhang
Department of Artificial Intelligence, Faculty of Computer Science and Systems Engineering, Kyushu Institute of Technology, Japan
Z
Zhihua Yao
Faculty of Economics and Business Administration, The University of Kitakyushu, Japan
K
Kouichi Sakurai
Department of Informatics, Faculty of Information Science and Electrical Engineering, Kyushu University, Japan
T
Takeshi Saitoh
Department of Artificial Intelligence, Faculty of Computer Science and Systems Engineering, Kyushu Institute of Technology, Japan