build hybrid codec

Design and implement a hybrid discrete–continuous codec that represents an input signal as a compressed sequence of discrete tokens plus a dimensionality‑reduced continuous residual and reconstructs the original with minimal information loss. Work covers methods to temporally compress token streams to reduce decoding steps, encode/quantize continuous residuals to preserve fine-grained characteristics lost by discretization, and decoder architectures that fuse token and residual pathways (e.g., modulation‑style fusion).

buildhybridcodec

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Discrete audio representations in speech language models often degrade downstream task performance due to information loss. To address this, this work proposes a hybrid discrete-continuous modeling approach that jointly represents speech using temporally compressed discrete tokens and dimensionality-reduced continuous residuals. The method introduces a novel encoder-decoder architecture incorporating fusion-focused modulation and a hybrid Transformer design, enabling autoregressive inference in the discrete domain while simultaneously leveraging non-autoregressive prediction and continuous residual upsampling. This approach achieves the first effective integration of discrete and continuous representations, substantially reducing the number of autoregressive steps while preserving speaker characteristics and fine-grained acoustic details. Experimental results demonstrate clear performance gains over purely discrete baselines.

discrete audio representationsinformation lossLarge Language Models

To address three key bottlenecks in aligning discrete acoustic codecs with speech large language models—training data scale mismatch, multi-codebook redundancy, and first-channel information overload—this paper proposes Mask Channel Residual Vector Quantization (MCRVQ). MCRVQ integrates a channel masking mechanism, an enhanced Fourier-based architecture, and training on 60K hours of speech data, enabling dynamic suppression of information density in the first RVQ channel and streamlined codebook hierarchy. Experiments demonstrate that MCRVQ significantly outperforms state-of-the-art codecs in audio compression. Moreover, in downstream speech-language modeling, it improves acoustic token generation quality and enhances robustness of text-to-speech alignment, achieving co-optimization between codec representations and language model capabilities.

Bridging gaps between discrete codecs and speech language modelsExcessive information in initial codebook channels hinders text-to-acoustic generationMultiple codebooks increase downstream speech model complexity

To address the challenges of subjective quality assessment and the inefficiency of purely data-driven models in speech/audio coding, this paper proposes a tightly integrated hybrid neural coding framework that synergistically combines model-driven and data-driven paradigms. Methodologically, it introduces a novel multi-level hybrid architecture that deeply couples psychoacoustic-weighted loss, customized time-frequency domain prediction (TF-Codec/MDCTNet), an LPCNet-based backbone, and a neural post-processing module, trained end-to-end via an autoencoder paradigm. The core contribution lies in systematically bridging the performance gap between classical signal modeling and end-to-end deep learning. Experimental results demonstrate that, at ultra-low bitrates of 1.6–3.2 kbps, the proposed method achieves a P.808 MOS gain of ≥0.5 over baselines, yielding subjective audio quality approaching that of wideband codecs, while increasing computational overhead by less than 15%.

Efficiency ImprovementNeural Voice and Audio CodingQuality Evaluation

LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec

Oct 21, 2024
YG
Yiwei Guo
🏛️ Shanghai Jiao Tong University

Discrete speech coding holds significant promise for language model–based speech generation, yet its adoption is hindered by high bitrates and strong entanglement between linguistic content and speaker identity. To address this, we propose LSCodec—a low-bitrate, speaker-decoupled discrete speech codec—introducing the first multi-stage unsupervised training framework. Our approach integrates a continuous information bottleneck, vector quantization, a discrete-token vocoder, and unsupervised speaker perturbation to achieve end-to-end mapping from continuous acoustic representations to a compact, speaker-disentangled discrete latent space. LSCodec employs only a single codebook with a reduced vocabulary size, yet surpasses baseline codecs in intelligibility and audio quality. Speaker conversion and speaker probing experiments demonstrate robust speaker disentanglement. Ablation studies validate the effectiveness of each component.

Decouples redundant timbre information from speechImproves intelligibility and audio quality efficientlyReduces high bitrate in discrete speech tokens

WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Aug 29, 2024
SJ
Shengpeng Ji
🏛️ Zhejiang University | Alibaba Group | Meta

To address the low compression efficiency and poor fidelity of discrete quantization for high-dimensional audio signals, this paper introduces WavTokenizer—the first efficient discrete encoder-decoder specifically designed for audio language modeling. It innovatively constructs a wide vector-quantized (VQ) codebook space, integrated with an extended-context Transformer, multi-scale GAN discriminators, and an inverse short-time Fourier transform (iSTFT)-based reconstruction architecture, enabling compression of 1-second 24 kHz audio into only 40–75 semantically rich tokens. WavTokenizer achieves state-of-the-art performance across speech, music, and general audio reconstruction: it attains a new industry-leading UTMOS score (+0.32), improves VQ codebook utilization by 37%, and significantly enhances both perceptual quality and semantic consistency. The model is lightweight, open-source, and natively compatible with downstream audio generation tasks.

Efficient audio signal compressionEnhancing subjective audio qualityOptimizing acoustic codec tokenizer

Latest Papers

What's happening recently
View more

This work addresses the inefficiency of diffusion models in lossy compression, where multi-step iterative sampling leads to slow encoding and decoding. To overcome this limitation, the authors integrate few-step generative models—such as Rectified Flow, Continuous Trajectory Matching (CTM), and MeanFlow—into the Reverse Channel Coding (RCC) framework, presenting the first probabilistic encoder-decoder formulation for these models that enables efficient compression without retraining. By leveraging velocity parameterization and denoising equivalence, they derive the posterior distribution required by RCC and further enhance CTM through EDM-inspired noise scheduling and local Gaussian approximation. Experiments demonstrate that the proposed approach substantially accelerates encoding and decoding on low-resolution benchmarks while improving perceptual quality of reconstructions at low bitrates.

diffusion modelsfew-step generative modelslossy compression

This work addresses the challenge of efficiently and disentangledly representing semantic and acoustic information in neural audio codecs by proposing HybridCodec, a novel codec that integrates a dual-stream architecture with self-supervised learning (SSL) representation distillation. During training, HybridCodec achieves strong disentanglement between semantic and acoustic features through semantic distillation, while eliminating the need for SSL models during inference. By unifying semantic distillation with a dual-stream structure, the method simultaneously ensures semantic specificity and high-fidelity reconstruction, significantly enhancing inference efficiency. Experimental results demonstrate that HybridCodec achieves superior semantic disentanglement and competitive reconstruction quality on in-domain tasks, exhibits robustness in cross-domain and zero-shot cross-lingual scenarios, and attains a threefold speedup in inference compared to existing dual-stream models.

acoustic featuresdual-stream architectureneural audio codec

This work addresses the suboptimal reconstruction in lossy compression caused by the mismatch between the encoder’s assumed source distribution and the true data distribution. To mitigate this issue without modifying the encoder, the authors propose a generative decompression framework that leverages prior knowledge of the true source distribution at the decoder. By employing Bayesian estimation and conditional expectation, the method achieves optimal reconstruction under the fixed encoder constraint. This study is the first to introduce Bayes-optimal decoding into mismatched compression scenarios and extends the approach to noisy channels and task-oriented compression. The framework integrates Gaussian source modeling, maximum a posteriori detection, and deep semantic classification, significantly narrowing the performance gap with jointly optimized benchmarks while enabling high-fidelity, adaptive reconstruction.

distribution mismatchgenerative decompressionlossy compression

This work addresses the challenges in autoregressive speech generation, where high frame rates or high-dimensional representations often lead to distributional drift and error accumulation, while low-dimensional representations compromise reconstruction fidelity. To overcome these limitations, the authors propose a jointly optimized framework featuring a low frame rate (8 Hz) yet high-dimensional (768-dim) continuous speech representation and a streaming generation architecture. Central to this approach is Locodec, a locally conditioned codec that enhances representational interpolability and coordinate identifiability, coupled with MP-ELD—a single-token autoregressive flow-matching mechanism incorporating multi-path routing and residual classifier-free guidance to effectively mitigate error propagation. Notably, the method achieves competitive word error rates (WER) and long-term stable, high-fidelity synthesis without relying on external SSL/ASR models, pretrained language models, or post-training, while maintaining high reconstruction quality and strong single-token predictability.

autoregressive speech generationerror accumulationhigh-dimensional tokens

Discrete visual generative models have long underperformed continuous counterparts due to limitations in codebook size and insufficient compression. This work proposes BAR (masked Bit AutoRegressive), a scalable autoregressive framework that decomposes discrete tokens into binary bit sequences and employs bit-wise masked modeling to enable efficient training and sampling with arbitrarily large codebooks. BAR overcomes the longstanding scalability and computational bottlenecks of discrete generative models, achieving a state-of-the-art 0.99 gFID on ImageNet-256—surpassing both leading continuous and discrete approaches—while significantly accelerating convergence and reducing sampling cost.

autoregressive modelingcodebook scalingdiscrete image generation