A Novel Semantic Compression Approach for Ultra-low Bandwidth Voice Communication

📅 2025-09-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing speech codecs do not explicitly disentangle semantic hierarchies, making it challenging to simultaneously preserve perceptual quality and downstream task performance at ultra-low bitrates (e.g., <1.5 kbps). To address this, we propose the first decoupled framework for semantic speech compression, introducing hierarchical semantic representations—explicitly separating and differentially encoding phonetic, prosodic, emotional, and speaker-related features—derived from generative speech models into the codec architecture. Leveraging a semantic communication paradigm and multi-granularity reconstruction, our method achieves or surpasses state-of-the-art performance of codecs such as EnCodec on automatic speech recognition, emotion analysis, and speaker verification, while operating at 2–4× lower bitrates. Crucially, intelligibility and naturalness are preserved. This work establishes a novel paradigm for ultra-low-bitrate semantic speech communication.

Technology Category

Natural Language Processing: SpeechMachine Learning: Deep Generative Models & AutoencodersCognitive Modeling & Cognitive Systems: Neural Spike Coding

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Vertical and domain-specific searchUser Modeling, Personalization and Recommendation: User modeling and simulation for interactive and conversational systems
📝 Abstract
While existing speech audio codecs designed for compression exploit limited forms of temporal redundancy and allow for multi-scale representations, they tend to represent all features of audio in the same way. In contrast, generative voice models designed for text-to-speech and voice transfer tasks have recently proved effective at factorizing audio signals into high-level semantic representations of fundamentally distinct features. In this paper, we leverage such representations in a novel semantic communications approach to achieve lower bitrates without sacrificing perceptual quality or suitability for specific downstream tasks. Our technique matches or outperforms existing audio codecs on transcription, sentiment analysis, and speaker verification when encoding at 2-4x lower bitrate -- notably surpassing Encodec in perceptual quality and speaker verification while using up to 4x less bitrate.
Problem

Research questions and friction points this paper is trying to address.

Achieving ultra-low bandwidth voice communication without quality loss
Leveraging semantic representations to reduce bitrates significantly
Outperforming existing codecs in perceptual quality and task performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic compression using generative voice models
Factorizing audio into distinct semantic representations
Achieving lower bitrates without sacrificing perceptual quality
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Ryan Collette
Systems & Technology Research, 600 W. Cummings Park, Woburn, MA 01801
R
Ross Greenwood
Systems & Technology Research, 600 W. Cummings Park, Woburn, MA 01801
S
Serena Nicoll
Systems & Technology Research, 600 W. Cummings Park, Woburn, MA 01801