🤖 AI Summary
Existing speech codecs do not explicitly disentangle semantic hierarchies, making it challenging to simultaneously preserve perceptual quality and downstream task performance at ultra-low bitrates (e.g., <1.5 kbps). To address this, we propose the first decoupled framework for semantic speech compression, introducing hierarchical semantic representations—explicitly separating and differentially encoding phonetic, prosodic, emotional, and speaker-related features—derived from generative speech models into the codec architecture. Leveraging a semantic communication paradigm and multi-granularity reconstruction, our method achieves or surpasses state-of-the-art performance of codecs such as EnCodec on automatic speech recognition, emotion analysis, and speaker verification, while operating at 2–4× lower bitrates. Crucially, intelligibility and naturalness are preserved. This work establishes a novel paradigm for ultra-low-bitrate semantic speech communication.
📝 Abstract
While existing speech audio codecs designed for compression exploit limited forms of temporal redundancy and allow for multi-scale representations, they tend to represent all features of audio in the same way. In contrast, generative voice models designed for text-to-speech and voice transfer tasks have recently proved effective at factorizing audio signals into high-level semantic representations of fundamentally distinct features. In this paper, we leverage such representations in a novel semantic communications approach to achieve lower bitrates without sacrificing perceptual quality or suitability for specific downstream tasks. Our technique matches or outperforms existing audio codecs on transcription, sentiment analysis, and speaker verification when encoding at 2-4x lower bitrate -- notably surpassing Encodec in perceptual quality and speaker verification while using up to 4x less bitrate.