🤖 AI Summary
Addressing the challenges of duration modeling and the trade-off between prosody and speaker identity in text-to-speech (TTS) for low-resource Indian languages, this work systematically investigates explicit duration prediction strategies—namely, speech completion-driven and duration prompt-enhanced approaches—within a non-autoregressive continuous normalizing flow (CNF) framework for single-speaker TTS. Experiments span multiple Indian languages and reveal a clear trade-off: completion-based prediction significantly improves intelligibility, whereas prompt-based strategies better preserve speaker timbre consistency. To our knowledge, this is the first study to empirically validate the effectiveness and interpretability of explicit duration modeling in low-resource multilingual TTS. The results establish a modular, controllable duration modeling paradigm—comprising decoupled, plug-in duration components—that advances resource-constrained TTS system design.
📝 Abstract
High-quality speech generation for low-resource languages, such as many Indian languages, remains a significant challenge due to limited data and diverse linguistic structures. Duration prediction is a critical component in many speech generation pipelines, playing a key role in modeling prosody and speech rhythm. While some recent generative approaches choose to omit explicit duration modeling, often at the cost of longer training times. We retain and explore this module to better understand its impact in the linguistically rich and data-scarce landscape of India. We train a non-autoregressive Continuous Normalizing Flow (CNF) based speech model using publicly available Indian language data and evaluate multiple duration prediction strategies for zero-shot, speaker-specific generation. Our comparative analysis on speech-infilling tasks reveals nuanced trade-offs: infilling based predictors improve intelligibility in some languages, while speaker-prompted predictors better preserve speaker characteristics in others. These findings inform the design and selection of duration strategies tailored to specific languages and tasks, underscoring the continued value of interpretable components like duration prediction in adapting advanced generative architectures to low-resource, multilingual settings.