BanglaBox: A Phonetically-Balanced Corpus and Data-Efficient Foundation-Model Adaptation for Bangla Text-to-Speech with Zero-Shot Voice Cloning

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of Bengali text-to-speech (TTS) corpora and the absence of open-source zero-shot voice cloning systems by proposing a comprehensive framework to efficiently adapt an English-pretrained autoregressive TTS model to Bengali. Methodologically, we design a hierarchical Jensen-Shannon divergence algorithm to construct a phoneme-balanced corpus, optimize the fine-tuning strategy through consistent tokenizer expansion coupled with a dual-loss objective incorporating prompt masking, and introduce text normalization processing. Experimental results demonstrate that the proposed system significantly outperforms baselines in both naturalness and speaker similarity, as validated by Wilcoxon signed-rank tests, while achieving few-shot performance using only one-seventh of the data volume. All code, models, and datasets have been made publicly available.
📝 Abstract
We present a recipe for adapting English-pretrained autoregressive TTS foundation models to underrepresented languages, demonstrated on Bangladeshi Bangla. Existing Bangla TTS corpora are small and single-speaker, and to our knowledge no open zero-shot voice-cloning system is available for the Bangladeshi register. We contribute a phonetically- and gender-balanced two-tier Bangladeshi Bangla corpus balanced via a tiered Jensen-Shannon divergence objective over conjunct clusters (juktakkhor), together with three fine-tuning changes: a merge-consistent tokenizer extension, Bangla text normalization, and a prompt-masked dual-loss objective. These changes preserve zero-shot cloning across the language switch. Our BanglaEval protocol applies Wilcoxon signed-rank tests with Bonferroni correction over native-speaker ratings. BanglaBox attains near-natural Naturalness, outperforms commercial and open-source baselines on Naturalness, Speaker Similarity, and Clarity, and reaches speaker similarity comparable to prior few-shot results while using approximately 7x less Bangla fine-tuning audio. We further validate the recipe beyond our own test split using a seven-category stress set of difficult real-world text, the public BnTTS evaluation benchmarks, and naturally occurring Bangla that no language model wrote. All artifacts, including the corpus, weights, code, and complete evaluation materials, are released publicly and unconditionally.
Problem

Research questions and friction points this paper is trying to address.

Bangla text-to-speech
zero-shot voice cloning
low-resource language adaptation
speech corpus
foundation model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Zero-Shot Voice Cloning
Foundation Model Adaptation
Phonetically-Balanced Corpus
Prompt-Masked Dual-Loss
Data-Efficient Fine-Tuning
🔎 Similar Papers
No similar papers found.