Implementation of Zero-shot Semantic Communication on Software Defined Radio

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of hardware validation and low transmission efficiency in zero-shot semantic communication by deploying a vision-language model (VLM) codec on a software-defined radio (SDR) platform. By integrating neural processing unit (NPU)-based edge acceleration with cosine similarity matching, this work provides the first empirical demonstration of zero-shot generalization to novel tasks over real wireless channels without retraining. Experimental results indicate that the proposed system achieves a ninefold bandwidth reduction compared to a JPEG baseline, attains 82% accuracy on CIFAR-10, and reduces encoding latency to 13.4 ms. These findings validate the feasibility of efficient, real-time semantic transmission under practical wireless conditions.
📝 Abstract
Semantic communication has recently gained traction for its ability to reduce the amount of data transmitted over a communication link by transmitting a task-oriented representation instead of the raw source. Zero-shot semantic communication sends a general embedding from a vision-language model (VLM), so the same transmitter can serve new classification tasks without retraining. Most evidence for this advantage, however, comes from numerical simulation. We implement zero-shot semantic communication on a software-defined radio platform: a Raspberry Pi drives a pair of Analog Devices Active Learning Module (ADALM)-Pluto transceivers, with an image encoder at the transmitter and a text encoder at the receiver, and determines the zero-shot classification results via cosine similarity. We compare two VLMs, CLIP and MobileCLIP, across various channel conditions, i.e., different signal-to-noise ratios (SNRs). We validate that the semantic link spends 9x fewer channel uses per image than a JPEG plus 16-ary quadrature amplitude modulation baseline and still reaches 82% accuracy on CIFAR-10 at 22.3 dB, where the baseline scores 0%. On the traffic sign recognition dataset (TSRD), MobileCLIP correctly classifies 98.3% of unseen images at the same SNR. Offloading the image encoder to a neural processing unit reduces encoding to 13.4 ms per image, 49x faster than a Raspberry Pi 4 CPU, placing the transmitter within a real-time budget. Our implementation is publicly available at https://github.com/thanhlexyz/zsscsdr.
Problem

Research questions and friction points this paper is trying to address.

Zero-shot semantic communication
Software Defined Radio
Vision-language model
Hardware implementation
Real-time processing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Zero-shot Semantic Communication
Software Defined Radio
Vision-Language Model
Real-time Edge Inference
Channel Efficiency
🔎 Similar Papers
No similar papers found.
Thanh Le
Thanh Le
Wireless System Laboratory - NICT
reinforcement learningwireless networks
A
Arif Dataesatu
Wireless Systems Laboratory, Wireless Networks Research Center, National Institute of Information and Communications Technology (NICT), Yokosuka, Kanagawa, Japan
H
Homare Murakami
Wireless Systems Laboratory, Wireless Networks Research Center, National Institute of Information and Communications Technology (NICT), Yokosuka, Kanagawa, Japan
T
Takeshi Matsumura
Wireless Systems Laboratory, Wireless Networks Research Center, National Institute of Information and Communications Technology (NICT), Yokosuka, Kanagawa, Japan