COT-TTS: Audio Context-Aware Text-to-Speech with Chain-of-Thought Reasoning

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决自然对话中语音风格自适应问题,提出COT-TTS系统,通过链式推理理解对话上下文并生成目标语音。
📝 Abstract
Recently, text-to-speech systems have made significant progress in speech expressiveness and controllability. However, the speaking style of generated speech typically relies on clear user-specified instructions. In natural conversations, speaking style should be naturally inferred from the preceding conversational context. Therefore, we propose COT-TTS, a context-aware, reasoning-based text-to-speech task. Given historical conversation audio, target text, and a reference speech, the system should comprehend the conversational context, infer an explicit intermediate reasoning, and finally synthesize the target speech with the specified timbre. To support this task, we constructed a large-scale bilingual conversational speech dataset comprising 9 million training samples, including a high-quality subset of 1 million samples. We further constructed a source-disjoint benchmark with 800 human-verified samples and established strong task-specific baselines. Additionally, we developed end-to-end autoregressive models with parameter sizes of 0.6B and 1.7B, generating emotion-labeled transcripts, editable speech style inferences, and speech tokens. Experimental results show that the proposed model achieves performance comparable to large-scale baseline systems with significantly fewer parameters. At the same time, the model performs well in terms of duration consistency and emotional consistency, and can generate appropriate emotional, stress, and rhythmic variations based on the conversational context. To facilitate future research, we will publicly release the data construction pipeline, dataset, trained models, and related resources. The demo page and additional resources are available at https://luckybian.github.io/COT-TTS
Problem

Research questions and friction points this paper is trying to address.

Text-to-Speech
Conversational Context
Chain-of-Thought Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

context-aware
chain-of-thought reasoning
bilingual conversational speech dataset
autoregressive models
emotional consistency
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Weizhen Bian
The Hong Kong University of Science and Technology, Hong Kong SAR, China
S
Sitong Cheng
The Hong Kong University of Science and Technology, Hong Kong SAR, China
R
Rongxiu Zhong
JIUTIAN Research, China Mobile, Beijing, China; The State Key Laboratory of Multimedia Information Processing, Peking University, Beijing, China
Jiahao Pan
Jiahao Pan
Hong Kong University of Science and Technology
Speech ProcessingSpeech EnhancmentMusic Generation
Liumeng Xue
Liumeng Xue
Hong Kong University of Science and Technology
Audio Speech and Language ProcessingSpeech Generation
Boyi Kang
Boyi Kang
The Hong Kong University of Science and Technology
Multimodal IntelligenceAudio Processing
S
Shilei Zhang
JIUTIAN Research, China Mobile, Beijing, China; The State Key Laboratory of Multimedia Information Processing, Peking University, Beijing, China
J
Jinglei Liu
China Mobile (Hong Kong) Innovation Research Institute, Hong Kong SAR, China
Y
Yue Wang
China Mobile (Hong Kong) Innovation Research Institute, Hong Kong SAR, China
Junlan Feng
Junlan Feng
Chief Scientist at China Mobile Research
Natural LanguageMachine LearningSpeech ProcessingData Mining
B
Bei Liu
The Hong Kong University of Science and Technology, Hong Kong SAR, China
Wei Xue
Wei Xue
Department of Applied Plant Science, Chonnam National University
Crop ecophysiology modellingclimate change