Score
Designs and builds curated speech datasets for training single‑speaker text‑to‑speech systems by collecting and transcribing utterances, recording and segmenting audio, and normalizing text and audio. Ensures phonetic and tonal coverage and prepares metadata, alignments, file formats, and audio trimming/normalization to meet TTS training requirements.
Training end-to-end text-to-speech (TTS) models with purely synthetic data remains underexplored, particularly regarding feasibility, robustness, and controllability compared to real speech data. Method: This study systematically evaluates FastSpeech 2- and VITS-based TTS models trained exclusively on synthetic speech, conducting ablation experiments by controlling textual richness, speaker diversity, environmental noise level, and speaking style. Evaluation integrates MOS, WER, CMOS, and subjective listening tests. Contribution/Results: To our knowledge, this is the first empirical demonstration that synthetic-data-only training achieves a MOS of 4.12—significantly surpassing the real-data baseline (3.78) at equivalent scale. The synthetic-trained models exhibit 27% higher robustness to accent and noise, and 31% improved cross-speaker generalization similarity. Key findings identify high text/speaker diversity and low environmental noise as primary drivers of robustness, while standard speaking style accelerates convergence. These results establish a theoretically grounded, cost-effective paradigm for controllable, high-quality TTS data curation.
High-quality text-to-speech (TTS) training is hindered by narrow domain coverage, licensing constraints, and insufficient scale of authentic speech data; meanwhile, large language model (LLM)-generated text suffers from low lexical diversity, existing text normalization tools lack robustness, and human recording is not scalable. To address these challenges, we propose SpeechWeave—the first end-to-end, automated multilingual synthetic speech data generation framework. It integrates prompt-optimized LLM-based text generation, a high-accuracy configurable text normalization module, and standardized TTS synthesis. SpeechWeave enables customizable, cross-lingual and cross-domain speech corpus construction, improving phonemic and linguistic diversity by 10–48%, achieving 97% text normalization accuracy, and producing highly consistent, TTS-optimized synthetic speech. The framework effectively alleviates the bottleneck imposed by real-world data limitations for large-scale TTS model training.
To address the escalating storage and annotation costs associated with rapidly growing TTS datasets, this paper proposes the first active learning framework for corpus construction in text-to-speech synthesis. Departing from conventional static, model-agnostic data collection paradigms, our approach establishes a closed-loop “sampling–modeling–feedback” pipeline: it dynamically selects high-informativeness text–speech pairs by jointly evaluating model uncertainty and sample diversity, followed by incremental model retraining. Experiments demonstrate that corpora constructed via our method significantly improve synthesized speech naturalness—yielding MOS gains of +0.3 to +0.5 under identical data scale—while achieving full baseline performance using only 60% of the data. This enables substantially more efficient and higher-quality data utilization for TTS development.
High-quality benchmark data for automatic naturalness assessment of Spanish text-to-speech (TTS) systems is lacking. Method: This paper introduces SPANAT—the first publicly available, diverse, and large-scale Spanish TTS naturalness evaluation dataset—comprising 4,326 audio samples from 52 TTS systems and human speech, subjectively annotated via Mean Opinion Score (MOS) by 92 participants following ITU-T P.807. We further propose two novel modeling paradigms leveraging self-supervised speech representations: (i) fine-tuning English-pretrained models (e.g., wav2vec 2.0), and (ii) attaching lightweight downstream networks to frozen encoders for cross-lingual transfer. Contribution/Results: Our best model achieves a mean absolute error (MAE) of 0.80 on the five-point MOS scale, validating both the dataset’s utility and the efficacy of the proposed approaches. SPANAT establishes a reproducible benchmark and advances methodology for naturalness assessment in low-resource languages.
Traditional TTS systems rely on studio-recorded, high-fidelity speech and thus suffer from poor generalization to real-world noisy environments; moreover, large-scale, in-the-wild speech-text paired datasets are scarce. To address this, we propose TITW—the first fully automated, large-scale in-the-wild TTS dataset—comprising two subsets: TITW-Easy and TITW-Hard. Constructed via metadata-driven crawling from VoxCeleb1, TITW integrates ASR-based transcription, DNSMOS-based quality assessment, and noise-aware data augmentation. Our approach eliminates reliance on manual annotation and clean speech, establishing a new benchmark for noise-robust TTS. TITW-Easy achieves UTMOS ≥ 3.0, while TITW-Hard provides challenging evaluation cases (UTMOS < 2.8). The dataset is publicly released, significantly advancing both research and practical deployment of in-the-wild TTS systems.
This work addresses the challenge faced by resource-constrained teams in developing high-performance text-to-speech (TTS) systems, which typically rely on massive proprietary datasets and complex multi-stage architectures. The authors propose a lightweight autoregressive TTS system featuring an extremely streamlined architecture, rigorous data engineering, and a novel Q-Former-based conditioning mechanism that effectively disentangles speaker identity from expressive style. This enables zero-shot voice cloning as well as synthesis of emotion, paralinguistic cues, and Chinese dialects. Trained exclusively on 200K hours of open-source data using a reproducible multi-stage preprocessing pipeline and cross-sample paired training, the system achieves a word error rate (WER) of 1.50% and character error rate (CER) of 0.87% on the Seed-TTS Eval benchmark for English and Chinese, respectively, with speaker similarity scores of 0.862 and 0.815—outperforming baselines trained on substantially larger datasets.
This work addresses the limitations of existing speech editing methods, which rely on task-specific training, incur high data costs, and struggle to preserve temporal consistency and speaker identity in unedited regions. The authors propose a training-free editing framework leveraging a pretrained autoregressive text-to-speech (TTS) model, enabling precise splicing between source and target speech through latent recomposition. To ensure natural transitions at edit boundaries without disrupting the generative manifold, they introduce Adaptive Weak Factor Guidance (AWFG). Additionally, they construct a new dataset, LibriSpeech-Edit, and propose a word-level dynamic time warping (WDTW) metric for evaluation. Experiments demonstrate that, compared to the strongest baseline, their method significantly improves temporal consistency in unedited segments and reduces word error rate by nearly 70%. When applied to a base TTS model, it achieves a 27% reduction in WDTW, setting a new state of the art in speaker identity preservation and temporal fidelity.
This study addresses the critical scarcity of multi-speaker conversational data in low-resource languages and niche domains, which severely limits the performance of conversational automatic speech recognition (ASR) systems. The authors propose a novel data synthesis approach that jointly leverages large language models (LLMs) and text-to-speech (TTS) systems: LLMs generate contextualized dialogues enriched with speaker metadata, and TTS synthesizes these texts while preserving speaker-specific attributes to produce speaker-aware simulated conversations. This work presents the first systematic validation of LLM–TTS–generated data for ASR training. Remarkably, a FastConformer-Large ASR model trained on only 67 hours of real data combined with 636 hours of synthetic data significantly outperforms a zero-shot baseline relying on 2,700 hours of real speech, demonstrating that high-quality synthetic data can effectively substitute for large volumes of authentic recordings.
This work addresses key challenges in multilingual, open-domain high-quality speech synthesis—namely zero-shot voice cloning, long-form generation stability, and fine-grained control over pronunciation and prosody—by proposing a streamlined yet scalable foundational speech generation model. Built upon discrete audio tokens and an autoregressive Transformer architecture, the approach introduces a causal Transformer-based audio tokenizer (MOSS-Audio-Tokenizer) and variable-bitrate residual vector quantization (RVQ) to construct a unified semantic-acoustic representation. It further incorporates a frame-wise local autoregressive module and a dual-generator mechanism, balancing modeling efficiency with deployment flexibility. The resulting model supports zero-shot voice cloning, phoneme- or pinyin-level pronunciation control, token-level duration adjustment, seamless code-switching, and low-latency initial-syllable output, enabling stable, high-fidelity, and speaker-consistent long-form speech synthesis across multiple languages.