Learnable Classifier-Free Guidance Null Embeddings for Enhanced Controllable Speech Synthesis

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入可学习的无条件嵌入替代固定空向量,增强可控语音合成的质量和鲁棒性,同时提供对说话人和文本引导的精细控制。
📝 Abstract
Classifier-free Guidance (CFG) is widely adopted in text-to-speech (TTS) systems to enhance generation quality and conditioning fidelity by interpolating between conditioned and unconditioned predictions. A common unconditional technique is to use an empty representation, in the form of a fixed null vector. In this work, we propose replacing this representation with a learnable unconditional embedding, optimized to represent a meaningful unconditional state. Objective and subjective evaluations demonstrate that learnable null embeddings consistently outperform fixed null embeddings across speaker similarity, speech stability, and expressiveness, while exhibiting greater robustness to larger guidance scales. We further show that learning a distinct unconditional embedding for each of the TTS conditioning modalities allows fine-grained control over speaker and text guidance, showcasing the trade-off between similarity and quality, and stability and expressiveness in the generated speech.
Problem

Research questions and friction points this paper is trying to address.

Classifier-free Guidance
Text-to-Speech
Controllable Speech Synthesis
Unconditional Embedding
Innovation

Methods, ideas, or system contributions that make the work stand out.

learnable null embeddings
classifier-free guidance
text-to-speech
unconditional embedding
fine-grained control
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Biel Tura Vecino
Biel Tura Vecino
Applied Scientist
Y
Yoach Lacombe
Cantina Labs
J
Julian Weber
Cantina Labs
Z
Zbigniew Łatka
Cantina Labs
Haitong Zhang
Haitong Zhang
Cantina Labs
L
Logan Hart
Cantina Labs
E
Eren Gölge
Cantina Labs