🤖 AI Summary
To address the poor robustness of image semantic transmission under fading channels and the difficulty of deploying semantic communication on IoT edge devices in 6G, this paper proposes SwinSIT—a lightweight, channel-adaptive semantic image transmission framework. Methodologically, it (1) designs an end-to-end semantic encoder-decoder based on Swin Transformer to explicitly model local-global image semantics; (2) introduces a novel SNR-feedback-driven collaborative enhancement mechanism between semantic maps and noise semantic maps, implemented via two-stage SNR-aware enhancement; (3) reuses an image denoising CNN for lightweight channel estimation and compensation (CEAC), augmented with SE attention to improve channel adaptability; and (4) employs joint pruning and quantization to drastically reduce model size. Experiments demonstrate that SwinSIT achieves significantly higher PSNR than conventional JSCC methods across diverse fading channels, reduces model size by over 70%, and maintains excellent reconstruction quality—enabling practical deployment on resource-constrained IoT edge devices.
📝 Abstract
Semantic communications (SCs) play a central role in shaping the future of the sixth generation (6G) wireless systems, which leverage rapid advances in deep learning (DL). In this regard, end-to-end optimized DL-based joint source-channel coding (JSCC) has been adopted to achieve SCs, particularly in image transmission. Utilizing vision transformers in the encoder/decoder design has enabled significant advancements in image semantic extraction, surpassing traditional convolutional neural networks (CNNs). In this paper, we propose a new JSCC paradigm for image transmission, namely Swin semantic image transmission (SwinSIT), based on the Swin transformer. The Swin transformer is employed to construct both the semantic encoder and decoder for efficient image semantic extraction and reconstruction. Inspired by the squeezing-and-excitation (SE) network, we introduce a signal-to-noise-ratio (SNR)-aware module that utilizes SNR feedback to adaptively perform a double-phase enhancement for the encoder-extracted semantic map and its noisy version at the decoder. Additionally, a CNN-based channel estimator and compensator (CEAC) module repurposes an image-denoising CNN to mitigate fading channel effects. To optimize deployment in resource-constrained IoT devices, a joint pruning and quantization scheme compresses the SwinSIT model. Simulations evaluate the SwinSIT performance against conventional benchmarks demonstrating its effectiveness. Moreover, the model's compressed version substantially reduces its size while maintaining favorable PSNR performance.