TwinViT-DeepJSCC: Adversarially Robust Semantic Image Communication

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of learning-based semantic communication systems to adversarial perturbation attacks by proposing a prevention-correction semantic image transceiver. The method employs a dual-ViT branch architecture with deep joint source-channel coding. At the receiver, it incorporates sensitivity-aware masking, confidence fusion, and a blind-estimated SNR-conditioned DDIM diffusion purification mechanism, enabling robust transmission under a fixed channel budget without requiring attack metadata. Experimental results demonstrate that the proposed framework achieves PSNR improvements of 9.5 dB under PGD attacks and 10.8 dB over Rayleigh fading channels, while increasing classification Top-1 accuracy by 38 percentage points compared to baselines.
📝 Abstract
Learning-based semantic communication is vulnerable to adversarial perturbations introduced before semantic encoding or over wireless channels. This paper proposes TwinViT-DeepJSCC, a preventive-corrective semantic image transceiver operating under a fixed channel-use budget. Two Vision Transformer (ViT)-based deep joint source-channel coding (DeepJSCC) branches learn complementary latent representations protected by sensitivity-aware masking. At the receiver, confidence-aware fusion, blind corruption-severity estimation, and signal-to-noise ratio (SNR)-severity-conditioned denoising diffusion implicit model (DDIM) purification mitigate residual corruption without requiring attack metadata. Experiments on the Canadian Institute for Advanced Research 100-class (CIFAR-100) dataset consider fast gradient sign method (FGSM), projected gradient descent (PGD), natural evolution strategies (NES), and Carlini-Wagner (CW) source-domain attacks, as well as random jamming and channel-aware adversarial waveforms over additive white Gaussian noise (AWGN) and block-flat Rayleigh fading. Under matched channel-use and attack budgets, TwinViT-DeepJSCC achieves maximum peak signal-to-noise ratio (PSNR) gains of approximately 9.5 dB under 20-step PGD and 10.8 dB under channel-aware waveform attacks over block-flat Rayleigh fading. Under PGD, it also improves Top-1 accuracy by up to approximately 38 percentage points over the undefended baseline and 13 percentage points over the strongest competing defense. Ablation results confirm the complementary contributions of the proposed transmitter- and receiver-side mechanisms.
Problem

Research questions and friction points this paper is trying to address.

semantic communication
adversarial robustness
DeepJSCC
wireless channel attacks
image transmission
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic Communication
DeepJSCC
Vision Transformer
Adversarial Robustness
Diffusion Model
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Maedeh Fallahreyhani
Department of Electrical Engineering, Tarbiat Modares University, Tehran, Iran
Paeiz Azmi
Paeiz Azmi
Department of Electrical Engineering, Tarbiat Modares University, Tehran, Iran
Nader Mokari
Nader Mokari
Department of Electrical Engineering, Tarbiat Modares University, Tehran, Iran
M
M. Reza Abedi
Department of Electrical Engineering, Tarbiat Modares University, Tehran, Iran
Melike Erol-Kantarci
Melike Erol-Kantarci
Canada Research Chair & Professor, University of Ottawa and Sr. Product Manager for AI RAN, Ericsson
AI-enabled wireless networksAIGenAI5G6GO-RANsmart gridAI\GenAI5G\6G\O-RAN
E
Eduard A. Jorswieck
Institute for Communications Technology, Technische Universität Braunschweig, Braunschweig, Germany