Technical report: Impact of Duration Prediction on Speaker-specific TTS for Indian Languages

📅 2025-07-22
📈 Citations: 0
Influential: 0
📄 PDF

career value

175K/year
🤖 AI Summary
Addressing the challenges of duration modeling and the trade-off between prosody and speaker identity in text-to-speech (TTS) for low-resource Indian languages, this work systematically investigates explicit duration prediction strategies—namely, speech completion-driven and duration prompt-enhanced approaches—within a non-autoregressive continuous normalizing flow (CNF) framework for single-speaker TTS. Experiments span multiple Indian languages and reveal a clear trade-off: completion-based prediction significantly improves intelligibility, whereas prompt-based strategies better preserve speaker timbre consistency. To our knowledge, this is the first study to empirically validate the effectiveness and interpretability of explicit duration modeling in low-resource multilingual TTS. The results establish a modular, controllable duration modeling paradigm—comprising decoupled, plug-in duration components—that advances resource-constrained TTS system design.

Technology Category

Application Category

📝 Abstract
High-quality speech generation for low-resource languages, such as many Indian languages, remains a significant challenge due to limited data and diverse linguistic structures. Duration prediction is a critical component in many speech generation pipelines, playing a key role in modeling prosody and speech rhythm. While some recent generative approaches choose to omit explicit duration modeling, often at the cost of longer training times. We retain and explore this module to better understand its impact in the linguistically rich and data-scarce landscape of India. We train a non-autoregressive Continuous Normalizing Flow (CNF) based speech model using publicly available Indian language data and evaluate multiple duration prediction strategies for zero-shot, speaker-specific generation. Our comparative analysis on speech-infilling tasks reveals nuanced trade-offs: infilling based predictors improve intelligibility in some languages, while speaker-prompted predictors better preserve speaker characteristics in others. These findings inform the design and selection of duration strategies tailored to specific languages and tasks, underscoring the continued value of interpretable components like duration prediction in adapting advanced generative architectures to low-resource, multilingual settings.
Problem

Research questions and friction points this paper is trying to address.

Improving speech generation for low-resource Indian languages
Evaluating duration prediction impact on prosody and rhythm
Optimizing speaker-specific TTS in multilingual, data-scarce contexts
Innovation

Methods, ideas, or system contributions that make the work stand out.

Non-autoregressive CNF model for Indian languages
Multiple duration prediction strategies evaluated
Speaker-prompted predictors preserve speaker characteristics
🔎 Similar Papers
No similar papers found.