Score
Designs and constructs datasets of synthetic speech by generating audio from text with TTS systems, producing aligned audio–transcript pairs and associated metadata. Builds augmented corpora that balance speaker and acoustic variability through TTS-based augmentation, and applies filtering and validation to ensure sample quality for training and evaluation.
Training end-to-end text-to-speech (TTS) models with purely synthetic data remains underexplored, particularly regarding feasibility, robustness, and controllability compared to real speech data. Method: This study systematically evaluates FastSpeech 2- and VITS-based TTS models trained exclusively on synthetic speech, conducting ablation experiments by controlling textual richness, speaker diversity, environmental noise level, and speaking style. Evaluation integrates MOS, WER, CMOS, and subjective listening tests. Contribution/Results: To our knowledge, this is the first empirical demonstration that synthetic-data-only training achieves a MOS of 4.12—significantly surpassing the real-data baseline (3.78) at equivalent scale. The synthetic-trained models exhibit 27% higher robustness to accent and noise, and 31% improved cross-speaker generalization similarity. Key findings identify high text/speaker diversity and low environmental noise as primary drivers of robustness, while standard speaking style accelerates convergence. These results establish a theoretically grounded, cost-effective paradigm for controllable, high-quality TTS data curation.
To address the escalating storage and annotation costs associated with rapidly growing TTS datasets, this paper proposes the first active learning framework for corpus construction in text-to-speech synthesis. Departing from conventional static, model-agnostic data collection paradigms, our approach establishes a closed-loop “sampling–modeling–feedback” pipeline: it dynamically selects high-informativeness text–speech pairs by jointly evaluating model uncertainty and sample diversity, followed by incremental model retraining. Experiments demonstrate that corpora constructed via our method significantly improve synthesized speech naturalness—yielding MOS gains of +0.3 to +0.5 under identical data scale—while achieving full baseline performance using only 60% of the data. This enables substantially more efficient and higher-quality data utilization for TTS development.
Current conversational TTS systems suffer from a scarcity of natural, interactive bilingual speech data, hindering effective modeling of authentic dialogue phenomena—such as overlapping speech, backchannel responses, and laughter. To address this, we introduce the first high-quality, bilingual (Chinese–English), full-duplex, spontaneous dialogue speech corpus (15 hours of multi-track recordings), covering everyday topics and genuine interactive behaviors. We propose an open-source dual-channel acquisition protocol and a fine-grained transcription annotation schema explicitly designed to capture overlapping utterances, nonverbal vocalizations, and feedback responses. Fine-tuning TTS models on this dataset yields statistically significant improvements over strong baselines in both objective metrics and subjective evaluations—particularly in speech naturalness and dialogue realism. This work establishes a foundational bilingual resource and methodological framework for advancing conversational speech synthesis.
High-quality text-to-speech (TTS) training is hindered by narrow domain coverage, licensing constraints, and insufficient scale of authentic speech data; meanwhile, large language model (LLM)-generated text suffers from low lexical diversity, existing text normalization tools lack robustness, and human recording is not scalable. To address these challenges, we propose SpeechWeave—the first end-to-end, automated multilingual synthetic speech data generation framework. It integrates prompt-optimized LLM-based text generation, a high-accuracy configurable text normalization module, and standardized TTS synthesis. SpeechWeave enables customizable, cross-lingual and cross-domain speech corpus construction, improving phonemic and linguistic diversity by 10–48%, achieving 97% text normalization accuracy, and producing highly consistent, TTS-optimized synthetic speech. The framework effectively alleviates the bottleneck imposed by real-world data limitations for large-scale TTS model training.
To address data scarcity in low-resource audio classification, this paper proposes a novel data augmentation framework integrating text-to-audio (T2A) diffusion modeling, preference optimization via proximal policy optimization (PPO), and large language model (LLM)-driven iterative caption generation. The method employs PPO to align synthesized audio with target acoustic characteristics, while the LLM dynamically generates and refines semantically diverse, high-quality captions—jointly enhancing acoustic fidelity and semantic richness. Distinct from prior work, this is the first study to holistically unify T2A diffusion, human-feedback-based preference optimization, and LLM-guided iterative captioning. Evaluated across 10 benchmark datasets under four low-resource settings, the framework achieves substantial performance gains (+0.1%–39%) using only a weakly supervised AudioSet-pretrained T2A model, consistently outperforming state-of-the-art baselines.
研究通过构建统一的基于音素的TTS增强管道及提出音素频率引导选择方法,解决自动语音识别中合成语音数据的有效利用问题。
This work addresses the degradation of speaker similarity in low-resource personalized text-to-speech (TTS) when naively augmenting training data with zero-shot TTS (ZS-TTS) synthesized speech. To mitigate this issue, the authors propose a lightweight domain-conditional training framework that distinguishes between real and synthetic speech through domain embeddings, without altering the base model architecture. Combined with an oversampling strategy for real data, this approach effectively preserves speaker characteristics during augmentation. Notably, it is the first to apply domain conditioning to ZS-TTS-based data augmentation. Experiments on LibriTTS and an internal dataset demonstrate that the method significantly outperforms naive augmentation under extremely limited target-speaker data, while maintaining high levels of naturalness, intelligibility, and speaker similarity.
This work addresses the challenge faced by resource-constrained teams in developing high-performance text-to-speech (TTS) systems, which typically rely on massive proprietary datasets and complex multi-stage architectures. The authors propose a lightweight autoregressive TTS system featuring an extremely streamlined architecture, rigorous data engineering, and a novel Q-Former-based conditioning mechanism that effectively disentangles speaker identity from expressive style. This enables zero-shot voice cloning as well as synthesis of emotion, paralinguistic cues, and Chinese dialects. Trained exclusively on 200K hours of open-source data using a reproducible multi-stage preprocessing pipeline and cross-sample paired training, the system achieves a word error rate (WER) of 1.50% and character error rate (CER) of 0.87% on the Seed-TTS Eval benchmark for English and Chinese, respectively, with speaker similarity scores of 0.862 and 0.815—outperforming baselines trained on substantially larger datasets.
In privacy-sensitive domains, the scarcity of real speech data and the distributional gap between synthetic and real speech hinder the effective use of synthetic data in automatic speech recognition (ASR). This work addresses this challenge within the SLAM-ASR framework by revealing, for the first time, that discriminative signals distinguishing real from synthetic speech in large language model (LLM) backbones are predominantly localized in early-to-mid layers. Leveraging this insight, the authors propose a synergistic strategy combining a layer selection module with room impulse response (RIR) augmentation. This approach substantially narrows the distributional gap, achieving performance on par with a full real-data baseline using only 25% of real speech (13.6 hours) and even surpassing it at higher proportions, thereby significantly reducing reliance on real speech data.
This study addresses the persistent challenges in low-resource speech synthesis—namely poor quality and weak generalization—stemming from scarce authentic corpora and orthographic diversity. To this end, we introduce OpenBibleTTS, the first large-scale multilingual text-to-speech (TTS) benchmark built upon authentic Bible texts and out-of-domain data, spanning 37 low-resource languages. We systematically evaluate state-of-the-art architectures, including Gemini-TTS and EveryVoice, in terms of intelligibility, naturalness, and cross-domain robustness. Our analysis reveals a trade-off between multilingual and monolingual approaches: Gemini-TTS achieves the highest subjective ratings across most languages, whereas monolingual EveryVoice demonstrates superior intelligibility for African languages. All data and models are publicly released to foster fair, reproducible research in low-resource TTS.