🤖 AI Summary
This work addresses the challenge of low-bitrate secure voice communication in IoT-aided non-terrestrial networks by proposing a physical-layer security scheme based on structured spectral compression. The transmitter employs segmented compressive sensing and quantization of Mel-spectra, integrated with an ARQ and forward error correction mechanism to ensure transmission reliability. At the receiver, lightweight speech reconstruction is achieved using a private dictionary matrix. By treating the compressive sensing dictionary as a cryptographic key, the method achieves high security: a mere 0.1% discrepancy in the dictionary reduces voiceprint similarity to 0.3. It operates at an ultralow bitrate of 3.9 kbps—lower than G.723—with only 12-bit memory footprint and O(n) linear time complexity, significantly outperforming existing approaches in terms of security, bandwidth efficiency, and computational overhead.
📝 Abstract
This paper focuses on the Low-Bitrate Secure Speech Communications based on the Structured Spectral Compression (LB-S2C2). Specifically, the Mel spectral matrix of the speech signal is first encoded at the transmitter side through compressive sensing based on waveform segmentation and data quantization. Then, the Automatic Repeat Request (ARQ) is combined with forward error correction to achieve reliable transmission of speech signals over wireless channels. Thirdly, the received signals are recovered as the speech at the receiver side. Finally, we conduct a series of simulation experiments for the performance evaluation of LB-S2C2. Our simulations reveal that the dictionary matrix used for the speech reconstruction is different from the one used for the high-order matrix sparsification by even only approximately 0.1%, and then the accurate speech recovery fails. It implies that the speech data can be securely transmitted when the dictionary matrix is preserved. More importantly, the LB-S2C2 exhibits a very high privacy protection capability with the average voiceprint similarity to be only 0.3, which is much lower than the 0.8 of the semantic speech communication scheme DeepSC-S, and even lower than the 0.33 of the latest speech communication scheme OFI-OFCNB. In addition, our simulations reveal that the proposed structured speech coding boasts a time complexity of merely O(n), and the proposed speech recovery scheme requires the 12-bit memory storage only, which outperforms the traditional encryption algorithms proposed for speech communications. In comparison with the conventional compression techniques, our spectral compression method renders the coding rate of only 3.9kbps, which is lower than the current lowest speech coding rate of 6.3kbps achieved by G.723.