tts dataset synthesis

Designs and constructs datasets of synthetic speech by generating audio from text with TTS systems, producing aligned audio–transcript pairs and associated metadata. Builds augmented corpora that balance speaker and acoustic variability through TTS-based augmentation, and applies filtering and validation to ensure sample quality for training and evaluation.

ttsdatasetsynthesis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.55
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Training end-to-end text-to-speech (TTS) models with purely synthetic data remains underexplored, particularly regarding feasibility, robustness, and controllability compared to real speech data. Method: This study systematically evaluates FastSpeech 2- and VITS-based TTS models trained exclusively on synthetic speech, conducting ablation experiments by controlling textual richness, speaker diversity, environmental noise level, and speaking style. Evaluation integrates MOS, WER, CMOS, and subjective listening tests. Contribution/Results: To our knowledge, this is the first empirical demonstration that synthetic-data-only training achieves a MOS of 4.12—significantly surpassing the real-data baseline (3.78) at equivalent scale. The synthetic-trained models exhibit 27% higher robustness to accent and noise, and 31% improved cross-speaker generalization similarity. Key findings identify high text/speaker diversity and low environmental noise as primary drivers of robustness, while standard speaking style accelerates convergence. These results establish a theoretically grounded, cost-effective paradigm for controllable, high-quality TTS data curation.

Assessing synthetic data's potential to outperform real data trainingExploring factors like speaker diversity and noise affecting model performanceInvestigating feasibility of purely synthetic data for TTS training

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection

Jul 11, 2025
KS
Kentaro Seki
🏛️ The University of Tokyo

To address the escalating storage and annotation costs associated with rapidly growing TTS datasets, this paper proposes the first active learning framework for corpus construction in text-to-speech synthesis. Departing from conventional static, model-agnostic data collection paradigms, our approach establishes a closed-loop “sampling–modeling–feedback” pipeline: it dynamically selects high-informativeness text–speech pairs by jointly evaluating model uncertainty and sample diversity, followed by incremental model retraining. Experiments demonstrate that corpora constructed via our method significantly improve synthesized speech naturalness—yielding MOS gains of +0.3 to +0.5 under identical data scale—while achieving full baseline performance using only 60% of the data. This enables substantially more efficient and higher-quality data utilization for TTS development.

Enhancing TTS quality through informative sample collectionImproving data efficiency for speech synthesis modelsReducing storage constraints in TTS dataset construction

Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech Synthesis

Sep 04, 2025
ZZ
Zhitong Zhou
🏛️ Beijing University of Civil Engineering and Architecture

Current conversational TTS systems suffer from a scarcity of natural, interactive bilingual speech data, hindering effective modeling of authentic dialogue phenomena—such as overlapping speech, backchannel responses, and laughter. To address this, we introduce the first high-quality, bilingual (Chinese–English), full-duplex, spontaneous dialogue speech corpus (15 hours of multi-track recordings), covering everyday topics and genuine interactive behaviors. We propose an open-source dual-channel acquisition protocol and a fine-grained transcription annotation schema explicitly designed to capture overlapping utterances, nonverbal vocalizations, and feedback responses. Fine-tuning TTS models on this dataset yields statistically significant improvements over strong baselines in both objective metrics and subjective evaluations—particularly in speech naturalness and dialogue realism. This work establishes a foundational bilingual resource and methodological framework for advancing conversational speech synthesis.

Addressing lack of realistic dual-speaker interaction dataCreating open-source conversational datasets for TTSEnhancing synthesized speech naturalness and interactivity

High-quality text-to-speech (TTS) training is hindered by narrow domain coverage, licensing constraints, and insufficient scale of authentic speech data; meanwhile, large language model (LLM)-generated text suffers from low lexical diversity, existing text normalization tools lack robustness, and human recording is not scalable. To address these challenges, we propose SpeechWeave—the first end-to-end, automated multilingual synthetic speech data generation framework. It integrates prompt-optimized LLM-based text generation, a high-accuracy configurable text normalization module, and standardized TTS synthesis. SpeechWeave enables customizable, cross-lingual and cross-domain speech corpus construction, improving phonemic and linguistic diversity by 10–48%, achieving 97% text normalization accuracy, and producing highly consistent, TTS-optimized synthetic speech. The framework effectively alleviates the bottleneck imposed by real-world data limitations for large-scale TTS model training.

Automating text normalization to improve data qualityGenerating diverse multilingual text for TTS trainingProducing scalable speaker-standardized synthetic speech audio

Synthio: Augmenting Small-Scale Audio Classification Datasets with Synthetic Data

Oct 02, 2024
SG
Sreyan Ghosh
🏛️ NVIDIA | University of Maryland

To address data scarcity in low-resource audio classification, this paper proposes a novel data augmentation framework integrating text-to-audio (T2A) diffusion modeling, preference optimization via proximal policy optimization (PPO), and large language model (LLM)-driven iterative caption generation. The method employs PPO to align synthesized audio with target acoustic characteristics, while the LLM dynamically generates and refines semantically diverse, high-quality captions—jointly enhancing acoustic fidelity and semantic richness. Distinct from prior work, this is the first study to holistically unify T2A diffusion, human-feedback-based preference optimization, and LLM-guided iterative captioning. Evaluated across 10 benchmark datasets under four low-resource settings, the framework achieves substantial performance gains (+0.1%–39%) using only a weakly supervised AudioSet-pretrained T2A model, consistently outperforming state-of-the-art baselines.

Ensuring acoustic consistency and diversity in synthetic dataGenerating synthetic audio that matches real-world diversityImproving audio classification with limited labeled data

Latest Papers

What's happening recently
View more

This work addresses the degradation of speaker similarity in low-resource personalized text-to-speech (TTS) when naively augmenting training data with zero-shot TTS (ZS-TTS) synthesized speech. To mitigate this issue, the authors propose a lightweight domain-conditional training framework that distinguishes between real and synthetic speech through domain embeddings, without altering the base model architecture. Combined with an oversampling strategy for real data, this approach effectively preserves speaker characteristics during augmentation. Notably, it is the first to apply domain conditioning to ZS-TTS-based data augmentation. Experiments on LibriTTS and an internal dataset demonstrate that the method significantly outperforms naive augmentation under extremely limited target-speaker data, while maintaining high levels of naturalness, intelligibility, and speaker similarity.

data augmentationlow-resourcepersonalized speech synthesis

This work addresses the challenge faced by resource-constrained teams in developing high-performance text-to-speech (TTS) systems, which typically rely on massive proprietary datasets and complex multi-stage architectures. The authors propose a lightweight autoregressive TTS system featuring an extremely streamlined architecture, rigorous data engineering, and a novel Q-Former-based conditioning mechanism that effectively disentangles speaker identity from expressive style. This enables zero-shot voice cloning as well as synthesis of emotion, paralinguistic cues, and Chinese dialects. Trained exclusively on 200K hours of open-source data using a reproducible multi-stage preprocessing pipeline and cross-sample paired training, the system achieves a word error rate (WER) of 1.50% and character error rate (CER) of 0.87% on the Seed-TTS Eval benchmark for English and Chinese, respectively, with speaker similarity scores of 0.862 and 0.815—outperforming baselines trained on substantially larger datasets.

barriers to entrydata efficiencyresource-constrained

In privacy-sensitive domains, the scarcity of real speech data and the distributional gap between synthetic and real speech hinder the effective use of synthetic data in automatic speech recognition (ASR). This work addresses this challenge within the SLAM-ASR framework by revealing, for the first time, that discriminative signals distinguishing real from synthetic speech in large language model (LLM) backbones are predominantly localized in early-to-mid layers. Leveraging this insight, the authors propose a synergistic strategy combining a layer selection module with room impulse response (RIR) augmentation. This approach substantially narrows the distributional gap, achieving performance on par with a full real-data baseline using only 25% of real speech (13.6 hours) and even surpassing it at higher proportions, thereby significantly reducing reliance on real speech data.

automatic speech recognitiondistributional gapLLM-based ASR

This study addresses the persistent challenges in low-resource speech synthesis—namely poor quality and weak generalization—stemming from scarce authentic corpora and orthographic diversity. To this end, we introduce OpenBibleTTS, the first large-scale multilingual text-to-speech (TTS) benchmark built upon authentic Bible texts and out-of-domain data, spanning 37 low-resource languages. We systematically evaluate state-of-the-art architectures, including Gemini-TTS and EveryVoice, in terms of intelligibility, naturalness, and cross-domain robustness. Our analysis reveals a trade-off between multilingual and monolingual approaches: Gemini-TTS achieves the highest subjective ratings across most languages, whereas monolingual EveryVoice demonstrates superior intelligibility for African languages. All data and models are publicly released to foster fair, reproducible research in low-resource TTS.

low-resource languagesmultilingual TTSspeech synthesis

Hot Scholars

AW

Alexander Waibel

Carnegie Mellon, KIT, Karlsruhe Institute of Technology, University of Karlsruhe
Machine LearningNeural NetworksSpeech TranslationMultimodal Interfaces
SW

Soo-Whan Chung

Naver Corporation
Speech Signal ProcessingMultimodal LearningGenerative Learning
BA

Bogdan Alexe

Faculty of Matematics and Computer Science, University of Bucharest
Computer VisionMachine LearningArtificial Intelligence
YL

Yu Liu

Tsinghua University
Computer Vision