Fill in the Gap! Combining Self-supervised Representation Learning with Neural Audio Synthesis for Speech Inpainting

📅 2024-05-30
🏛️ arXiv.org
📈 Citations: 3
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether the self-supervised speech representation model HuBERT can directly support speech inpainting—reconstructing missing or corrupted speech segments—without task-specific fine-tuning. We propose an encoder–decoder collaborative modeling framework: a frozen or fine-tuned HuBERT encoder is coupled with a HiFi-GAN vocoder decoder to jointly model context-aware waveform generation. Our key contribution is the first explicit alignment of self-supervised pretraining objectives with speech inpainting, supporting both known and unknown mask locations, as well as single- and multi-speaker scenarios. Experiments demonstrate that fine-tuning HuBERT achieves precise reconstruction of up to 400-ms segments in single-speaker settings; in multi-speaker settings, freezing HuBERT while optimizing HiFi-GAN significantly improves naturalness and intelligibility. Both objective metrics (e.g., PESQ, STOI) and subjective listening evaluations confirm the effectiveness of our approach.

Technology Category

Natural Language Processing: SpeechMachine Learning: Unsupervised & Self-Supervised LearningComputer Vision: Diffusion Models for Vision

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Most speech self-supervised learning (SSL) models are trained with a pretext task which consists in predicting missing parts of the input signal, either future segments (causal prediction) or segments masked anywhere within the input (non-causal prediction). Learned speech representations can then be efficiently transferred to downstream tasks (e.g., automatic speech or speaker recognition). In the present study, we investigate the use of a speech SSL model for speech inpainting, that is reconstructing a missing portion of a speech signal from its surrounding context, i.e., fulfilling a downstream task that is very similar to the pretext task. To that purpose, we combine an SSL encoder, namely HuBERT, with a neural vocoder, namely HiFiGAN, playing the role of a decoder. In particular, we propose two solutions to match the HuBERT output with the HiFiGAN input, by freezing one and fine-tuning the other, and vice versa. Performance of both approaches was assessed in single- and multi-speaker settings, for both informed and blind inpainting configurations (i.e., the position of the mask is known or unknown, respectively), with different objective metrics and a perceptual evaluation. Performances show that if both solutions allow to correctly reconstruct signal portions up to the size of 200ms (and even 400ms in some cases), fine-tuning the SSL encoder provides a more accurate signal reconstruction in the single-speaker setting case, while freezing it (and training the neural vocoder instead) is a better strategy when dealing with multi-speaker data.
Problem

Research questions and friction points this paper is trying to address.

Investigates SSL-trained encoders for speech inpainting without extra training
Compares SSL-based methods to supervised fine-tuning for speech reconstruction
Evaluates inpainting under varied conditions including unseen speakers and noise
Innovation

Methods, ideas, or system contributions that make the work stand out.

Using SSL-trained encoder without additional training
Adding decoder to generate waveform for inpainting
Fine-tuning encoder or decoder for different scenarios
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.