Towards Robust Generative Speech Enhancement Using Vector Quantisation-Based Neural Audio Codec

📅 2026-06-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited robustness of generative speech enhancement under complex noise conditions by proposing two latent-space modeling frameworks: continuous (cNAC-SE) and discrete (dNAC-SE). The core innovation lies in introducing a clean-prior-based vector quantization (VQ) regularization, which substantially improves model robustness. Crucially, this performance gain stems from the structural constraints imposed by VQ on latent representations rather than the discreteness of tokens per se, enabling effective transfer to continuous modeling paradigms. Systematic analysis reveals key mechanistic differences between the two frameworks. Experimental results demonstrate that the fully fine-tuned cNAC-SE consistently outperforms its dNAC-SE counterparts across diverse noise conditions, achieving state-of-the-art performance among generative approaches on the DNS-MOS benchmark.
📝 Abstract
This work investigates modelling strategies in continuous and discrete latent spaces in the vector quantisation (VQ)-based neural audio codec (NAC) speech enhancement (SE), along with the role of VQ regularisation. We propose cNAC-SE and dNAC-SE frameworks that predict continuous representations and discrete tokens in latent space, respectively. Theoretical analysis and visualisations in latent space are performed to exhibit their inherent modelling mechanisms. Experimental results show that the fully fine-tuned cNAC-SE model consistently outperforms all dNAC-SE variants across diverse test conditions and achieves leading performance among established generative approaches in DNS-MOS metrics. Comparison with the discriminative counterpart shows that VQ enhances robustness through an intrinsic effect of clean-prior-constrained regularisation, independent of discrete token processing. This highlights the transferable value of VQ regularisation to other continuous modelling methods.
Problem

Research questions and friction points this paper is trying to address.

speech enhancement
robustness
generative models
vector quantisation
neural audio codec
Innovation

Methods, ideas, or system contributions that make the work stand out.

vector quantisation
neural audio codec
speech enhancement
latent space modelling
VQ regularisation
🔎 Similar Papers
No similar papers found.