Score
Designs, implements, and evaluates systems that convert written text into spoken audio, encompassing text normalization and phoneme conversion, prosody and duration modeling, voice/timbre representation, and waveform generation (including neural vocoders), with attention to intelligibility, naturalness, expressiveness, and latency.
Existing text-to-speech (TTS) and text-to-music (TTM) systems rely separately on reference audio or expert-annotated heterogeneous conditions, lacking a unified, natural-language-driven framework for fine-grained, multi-attribute control—and have long been modeled in isolation. Method: We propose the first natural-language-instruction-based unified speech and music generation framework, enabling cross-modal control over timbre, emotion, style, language, instrumentation, tempo, and more. Our approach employs a joint-single diffusion Transformer architecture with standardized instruction-phoneme inputs, trained end-to-end on 50K hours of speech and 20K hours of music data to achieve cross-modal alignment and multi-task learning. Contribution/Results: The framework generates expressive bilingual (Chinese/English) speech, music, and spoken dialogue. Experiments demonstrate state-of-the-art performance across standard metrics, validating the effectiveness and generalizability of instruction-driven unified generation.
This work proposes an end-to-end speech generation approach driven by natural language instructions, addressing the limited expressiveness and diversity of current text-to-speech models that predominantly rely on studio-recorded data. By leveraging large-scale movie dialogue corpora and integrating open-source instruction-following architectures with in-the-wild speech, the method enables direct synthesis of highly realistic voices from free-form textual descriptions specifying character traits, personality, and emotional states. Subjective evaluations demonstrate that the proposed model significantly outperforms existing voice design systems in overall audio quality, fidelity to user instructions, and naturalness, marking a substantial advance toward human-like vocal expressivity in synthetic speech.
High-quality text-to-speech (TTS) training is hindered by narrow domain coverage, licensing constraints, and insufficient scale of authentic speech data; meanwhile, large language model (LLM)-generated text suffers from low lexical diversity, existing text normalization tools lack robustness, and human recording is not scalable. To address these challenges, we propose SpeechWeave—the first end-to-end, automated multilingual synthetic speech data generation framework. It integrates prompt-optimized LLM-based text generation, a high-accuracy configurable text normalization module, and standardized TTS synthesis. SpeechWeave enables customizable, cross-lingual and cross-domain speech corpus construction, improving phonemic and linguistic diversity by 10–48%, achieving 97% text normalization accuracy, and producing highly consistent, TTS-optimized synthetic speech. The framework effectively alleviates the bottleneck imposed by real-world data limitations for large-scale TTS model training.
Traditional neural speech codecs struggle to disentangle linguistic content, speaker identity, and prosody, often resulting in poor prosody preservation during voice conversion. This work proposes a prosody-oriented codec that models prosody as a conditional residual guided by textual and speaker embeddings, while capturing prosodic variations unexplained by content or speaker through discrete bottleneck representations. By integrating low-frequency Mel-band modeling and training on same-speaker paired data, the method effectively enhances prosody retention. Experimental results demonstrate that the proposed approach significantly improves prosody transfer in voice conversion tasks and substantially reduces source speaker timbre leakage.
Existing text-to-audio (TTA) models rely on coarse-grained textual descriptions, limiting fine-grained control over both content and style; incorporating frame-level conditioning or dedicated control networks introduces architectural complexity and degrades generalization. This paper proposes the first purely natural-language-driven fine-grained TTA framework. First, we design an automated fine-grained data simulation pipeline to synthesize high-quality text–audio pairs with precise semantic–acoustic alignment. Second, we introduce a streaming diffusion Transformer architecture that integrates cross-attention mechanisms to faithfully align linguistic semantics with temporal audio features. Crucially, our method requires no frame-level annotations or auxiliary control modules. Despite its compact model size and faster inference speed, it achieves superior audio fidelity and fine-grained controllability compared to state-of-the-art approaches.
This work addresses the limitations in current non-human voice conversion research, which has been hindered by the absence of publicly available, structured datasets and an overreliance on natural human speech. To bridge this gap, we present the first systematically constructed and open-sourced dataset encompassing both human and animal vocalizations, augmented with professionally designed audio effects. The dataset explicitly disentangles timbre and style dimensions and incorporates a structured train-test split strategy, enabling rigorous evaluation of models’ controllable generalization capabilities across both seen and unseen timbres and styles. By providing a high-quality benchmark and reproducible foundation, this resource significantly advances research in non-human voice conversion.
This study addresses the communication needs of individuals with speech impairments by proposing a high-fidelity speech reconstruction method based on intracranial electroencephalography (iEEG). The approach systematically extracts prosodic features—such as intonation, pitch, and rhythm—from iEEG signals and introduces a novel Transformer-based encoder architecture that explicitly incorporates prosodic information. This is the first work to achieve prosody-driven natural speech synthesis in brain-to-speech tasks. Experimental results demonstrate that the proposed method significantly outperforms baseline models, including Griffin–Lim and CNN-based approaches, both in objective metrics and subjective listening evaluations. The synthesized speech exhibits markedly improved intelligibility and expressiveness, highlighting the critical role of neural prosody encoding in reconstructing naturalistic vocal output from brain activity.
This study addresses the lack of comprehensive evaluation frameworks for cross-lingual text-to-speech (TTS) systems in low-resource languages that jointly assess perceptual quality, speaker similarity, and acoustic fidelity. The authors propose a reproducible multi-metric benchmark integrating MUSHRA/ABX subjective listening tests, Resemblyzer-based speaker similarity scores, and objective measures such as Mel-cepstral distortion (MCD) and F0 RMSE. For the first time, four state-of-the-art TTS systems are rigorously evaluated across four distinct speech domains—formal, conversational, literary, and emotional—in a low-resource setting. Results reveal that emotional speech synthesis poses the greatest challenge (average MCD: 12.03 dB), while conversational speech achieves the highest acoustic fidelity, with significant performance variations observed across systems and domains. The complete evaluation toolkit and dataset are publicly released to advance standardized assessment in low-resource TTS research.
This work addresses the ambiguity in labeling resynthesized audio generated by neural audio codecs—a class of models that combine compression and synthesis capabilities—within the context of voice spoofing detection. The study presents the first systematic analysis of this labeling challenge, introducing an extended version of the ASVspoof 5 dataset and proposing multiple annotation strategies tailored to resynthesized audio. A unified evaluation framework is designed to assess the impact of different labeling approaches on anti-spoofing systems, leveraging resynthesis techniques that integrate neural codecs with vocoders. Experimental results demonstrate that the choice of annotation strategy significantly influences detection performance, offering critical insights for future dataset construction and evaluation protocols in audio deepfake detection research.