Narrowband Voice Communication Using Streaming Neural Compression

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决低功耗设备上的窄带语音通信问题,提出了一种轻量级神经音频编解码器TinyCall,并通过改进的量化方法和训练策略实现高效部署。
📝 Abstract
Low-bitrate speech communication on resource-constrained edge devices remains challenging due to stringent computational, memory, and bandwidth constraints. We present TinyCall, a lightweight neural audio codec designed for real-time speech communication on low-power platforms such as the ESP32 microcontroller and Raspberry Pi. The proposed system targets emergency communication and other bandwidth-limited scenarios while preserving speech intelligibility, speaker identity, and vocal expressiveness. To enable efficient deployment, we propose a minimal neural audio codec architecture together with a framework for converting a causally trained codec into a truly streamable codec through pseudo-lookahead decoding and decoder-input caching. We further replace conventional residual vector quantization (RVQ) with Residual Finite Scalar Quantization (RFSQ) to reduce inference complexity on edge processors and employ a progressive three-stage training strategy for stable optimization under latent quantization. An MFCC-based perceptual loss encourages preservation of speaker characteristics, including harmonic structure and vocal timbre. Experimental results demonstrate real-time operation on a Raspberry Pi 3 while achieving intelligible speech reconstruction at bitrates as low as 2.3 kbps. The proposed approach demonstrates that practical neural speech communication is feasible on highly resource-constrained edge devices.
Problem

Research questions and friction points this paper is trying to address.

Low-bitrate speech communication
Resource-constrained edge devices
Computational constraints
Memory constraints
Bandwidth constraints
Innovation

Methods, ideas, or system contributions that make the work stand out.

Neural Audio Codec
Pseudo-lookahead Decoding
Residual Finite Scalar Quantization (RFSQ)
Perceptual Loss
🔎 Similar Papers
No similar papers found.