A Novel Semantic Compression Approach for Ultra-low Bandwidth Voice Communication
Existing speech codecs do not explicitly disentangle semantic hierarchies, making it challenging to simultaneously preserve perceptual quality and downstream task performance at ultra-low bitrates (e.g., <1.5 kbps). To address this, we propose the first decoupled framework for semantic speech compression, introducing hierarchical semantic representations—explicitly separating and differentially encoding phonetic, prosodic, emotional, and speaker-related features—derived from generative speech models into the codec architecture. Leveraging a semantic communication paradigm and multi-granularity reconstruction, our method achieves or surpasses state-of-the-art performance of codecs such as EnCodec on automatic speech recognition, emotion analysis, and speaker verification, while operating at 2–4× lower bitrates. Crucially, intelligibility and naturalness are preserved. This work establishes a novel paradigm for ultra-low-bitrate semantic speech communication.