🤖 AI Summary
This study addresses the scarcity of Bengali text-to-speech (TTS) corpora and the absence of open-source zero-shot voice cloning systems by proposing a comprehensive framework to efficiently adapt an English-pretrained autoregressive TTS model to Bengali. Methodologically, we design a hierarchical Jensen-Shannon divergence algorithm to construct a phoneme-balanced corpus, optimize the fine-tuning strategy through consistent tokenizer expansion coupled with a dual-loss objective incorporating prompt masking, and introduce text normalization processing. Experimental results demonstrate that the proposed system significantly outperforms baselines in both naturalness and speaker similarity, as validated by Wilcoxon signed-rank tests, while achieving few-shot performance using only one-seventh of the data volume. All code, models, and datasets have been made publicly available.
📝 Abstract
We present a recipe for adapting English-pretrained autoregressive TTS foundation models to underrepresented languages, demonstrated on Bangladeshi Bangla. Existing Bangla TTS corpora are small and single-speaker, and to our knowledge no open zero-shot voice-cloning system is available for the Bangladeshi register. We contribute a phonetically- and gender-balanced two-tier Bangladeshi Bangla corpus balanced via a tiered Jensen-Shannon divergence objective over conjunct clusters (juktakkhor), together with three fine-tuning changes: a merge-consistent tokenizer extension, Bangla text normalization, and a prompt-masked dual-loss objective. These changes preserve zero-shot cloning across the language switch. Our BanglaEval protocol applies Wilcoxon signed-rank tests with Bonferroni correction over native-speaker ratings. BanglaBox attains near-natural Naturalness, outperforms commercial and open-source baselines on Naturalness, Speaker Similarity, and Clarity, and reaches speaker similarity comparable to prior few-shot results while using approximately 7x less Bangla fine-tuning audio. We further validate the recipe beyond our own test split using a seven-category stress set of difficult real-world text, the public BnTTS evaluation benchmarks, and naturally occurring Bangla that no language model wrote. All artifacts, including the corpus, weights, code, and complete evaluation materials, are released publicly and unconditionally.