text-to-speech

Designs, implements, and evaluates systems that convert written text into spoken audio, encompassing text normalization and phoneme conversion, prosody and duration modeling, voice/timbre representation, and waveform generation (including neural vocoders), with attention to intelligibility, naturalness, expressiveness, and latency.

text-to-speech

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.96
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$229K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

InstructAudio: Unified speech and music generation with natural language instruction

Nov 23, 2025
CQ
Chunyu Qiang
🏛️ Tianjin University | Kuaishou Technology | Institute of Automation, Chinese Academy of Sciences

Existing text-to-speech (TTS) and text-to-music (TTM) systems rely separately on reference audio or expert-annotated heterogeneous conditions, lacking a unified, natural-language-driven framework for fine-grained, multi-attribute control—and have long been modeled in isolation. Method: We propose the first natural-language-instruction-based unified speech and music generation framework, enabling cross-modal control over timbre, emotion, style, language, instrumentation, tempo, and more. Our approach employs a joint-single diffusion Transformer architecture with standardized instruction-phoneme inputs, trained end-to-end on 50K hours of speech and 20K hours of music data to achieve cross-modal alignment and multi-task learning. Contribution/Results: The framework generates expressive bilingual (Chinese/English) speech, music, and spoken dialogue. Experiments demonstrate state-of-the-art performance across standard metrics, validating the effectiveness and generalizability of instruction-driven unified generation.

No unified framework exists for joint speech and music generation via instructionsTTM systems require expert annotations and have limited input conditioningTTS systems lack natural language control over timbre and dialogue generation

This work proposes an end-to-end speech generation approach driven by natural language instructions, addressing the limited expressiveness and diversity of current text-to-speech models that predominantly rely on studio-recorded data. By leveraging large-scale movie dialogue corpora and integrating open-source instruction-following architectures with in-the-wild speech, the method enables direct synthesis of highly realistic voices from free-form textual descriptions specifying character traits, personality, and emotional states. Subjective evaluations demonstrate that the proposed model significantly outperforms existing voice design systems in overall audio quality, fidelity to user instructions, and naturalness, marking a substantial advance toward human-like vocal expressivity in synthetic speech.

expressive speechnatural language descriptionstext-to-speech

High-quality text-to-speech (TTS) training is hindered by narrow domain coverage, licensing constraints, and insufficient scale of authentic speech data; meanwhile, large language model (LLM)-generated text suffers from low lexical diversity, existing text normalization tools lack robustness, and human recording is not scalable. To address these challenges, we propose SpeechWeave—the first end-to-end, automated multilingual synthetic speech data generation framework. It integrates prompt-optimized LLM-based text generation, a high-accuracy configurable text normalization module, and standardized TTS synthesis. SpeechWeave enables customizable, cross-lingual and cross-domain speech corpus construction, improving phonemic and linguistic diversity by 10–48%, achieving 97% text normalization accuracy, and producing highly consistent, TTS-optimized synthetic speech. The framework effectively alleviates the bottleneck imposed by real-world data limitations for large-scale TTS model training.

Automating text normalization to improve data qualityGenerating diverse multilingual text for TTS trainingProducing scalable speaker-standardized synthetic speech audio

Traditional neural speech codecs struggle to disentangle linguistic content, speaker identity, and prosody, often resulting in poor prosody preservation during voice conversion. This work proposes a prosody-oriented codec that models prosody as a conditional residual guided by textual and speaker embeddings, while capturing prosodic variations unexplained by content or speaker through discrete bottleneck representations. By integrating low-frequency Mel-band modeling and training on same-speaker paired data, the method effectively enhances prosody retention. Experimental results demonstrate that the proposed approach significantly improves prosody transfer in voice conversion tasks and substantially reduces source speaker timbre leakage.

disentanglementneural speech representationprosody

AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions

Sep 19, 2024
YW
Yuanyuan Wang
🏛️ The Chinese University of Hong Kong | Tencent AI Lab | Tsinghua University

Existing text-to-audio (TTA) models rely on coarse-grained textual descriptions, limiting fine-grained control over both content and style; incorporating frame-level conditioning or dedicated control networks introduces architectural complexity and degrades generalization. This paper proposes the first purely natural-language-driven fine-grained TTA framework. First, we design an automated fine-grained data simulation pipeline to synthesize high-quality text–audio pairs with precise semantic–acoustic alignment. Second, we introduce a streaming diffusion Transformer architecture that integrates cross-attention mechanisms to faithfully align linguistic semantics with temporal audio features. Crucially, our method requires no frame-level annotations or auxiliary control modules. Despite its compact model size and faster inference speed, it achieves superior audio fidelity and fine-grained controllability compared to state-of-the-art approaches.

Addresses data scarcity with automated fine-grained text-audio pairingEnables fine-grained audio generation using natural language descriptionsImproves text-to-audio models without complex control networks

Latest Papers

What's happening recently
View more

This work addresses the limitations in current non-human voice conversion research, which has been hindered by the absence of publicly available, structured datasets and an overreliance on natural human speech. To bridge this gap, we present the first systematically constructed and open-sourced dataset encompassing both human and animal vocalizations, augmented with professionally designed audio effects. The dataset explicitly disentangles timbre and style dimensions and incorporates a structured train-test split strategy, enabling rigorous evaluation of models’ controllable generalization capabilities across both seen and unseen timbres and styles. By providing a high-quality benchmark and reproducible foundation, this resource significantly advances research in non-human voice conversion.

audio datasetdesigned vocalizationsnon-human voices

This study addresses the communication needs of individuals with speech impairments by proposing a high-fidelity speech reconstruction method based on intracranial electroencephalography (iEEG). The approach systematically extracts prosodic features—such as intonation, pitch, and rhythm—from iEEG signals and introduces a novel Transformer-based encoder architecture that explicitly incorporates prosodic information. This is the first work to achieve prosody-driven natural speech synthesis in brain-to-speech tasks. Experimental results demonstrate that the proposed method significantly outperforms baseline models, including Griffin–Lim and CNN-based approaches, both in objective metrics and subjective listening evaluations. The synthesized speech exhibits markedly improved intelligibility and expressiveness, highlighting the critical role of neural prosody encoding in reconstructing naturalistic vocal output from brain activity.

brain-to-speechiEEGneuroprosthetics

This study addresses the lack of comprehensive evaluation frameworks for cross-lingual text-to-speech (TTS) systems in low-resource languages that jointly assess perceptual quality, speaker similarity, and acoustic fidelity. The authors propose a reproducible multi-metric benchmark integrating MUSHRA/ABX subjective listening tests, Resemblyzer-based speaker similarity scores, and objective measures such as Mel-cepstral distortion (MCD) and F0 RMSE. For the first time, four state-of-the-art TTS systems are rigorously evaluated across four distinct speech domains—formal, conversational, literary, and emotional—in a low-resource setting. Results reveal that emotional speech synthesis poses the greatest challenge (average MCD: 12.03 dB), while conversational speech achieves the highest acoustic fidelity, with significant performance variations observed across systems and domains. The complete evaluation toolkit and dataset are publicly released to advance standardized assessment in low-resource TTS research.

acoustic fidelitydomain-specific analysislow-resource languages

This work addresses the ambiguity in labeling resynthesized audio generated by neural audio codecs—a class of models that combine compression and synthesis capabilities—within the context of voice spoofing detection. The study presents the first systematic analysis of this labeling challenge, introducing an extended version of the ASVspoof 5 dataset and proposing multiple annotation strategies tailored to resynthesized audio. A unified evaluation framework is designed to assess the impact of different labeling approaches on anti-spoofing systems, leveraging resynthesis techniques that integrate neural codecs with vocoders. Experimental results demonstrate that the choice of annotation strategy significantly influences detection performance, offering critical insights for future dataset construction and evaluation protocols in audio deepfake detection research.

audio deepfake detectionlabeling ambiguityneural audio codecs

Hot Scholars

HQ

Huamin Qu

Chair Professor, Hong Kong University of Science and Technology
Data visualizationHuman-Computer InteractionExplainable AIE-Learning
JC

Jiannong Cao

IEEE Fellow; Chair Professor, Hong Kong Polytechnic University
Distributed computingMobile and pervasive computingWireless sensor networksCloud computing
ZW

Zhiyuan Wen

The Hong Kong Polytechnic University
NLP
XJ

Xiaogang Jin

Professor of the State Key Lab of CAD&CG, Zhejiang University
Computer AnimationComputer GraphicsVirtual RealityDigital Fashion