Assessing the Ability of Neural TTS Systems to Model Consonant-Induced F0 Perturbation

📅 2026-03-22
📈 Citations: 0
Influential: 0
📄 PDF

career value

186K/year
🤖 AI Summary
It remains unclear whether current neural text-to-speech (TTS) systems can accurately model fine-grained segmental prosodic phenomena—specifically, consonant-induced fundamental frequency (F0) perturbations—and how well they generalize to low-frequency words. This work proposes a linguistically motivated, segment-level prosody probing framework, training Tacotron 2 and FastSpeech 2 models on the LJ Speech corpus and evaluating their F0 modeling capabilities through large-scale multi-system comparisons and lexical frequency-stratified analyses. Results indicate that while both systems perform well on high-frequency words, they exhibit limited generalization on low-frequency items, suggesting reliance on lexical memorization rather than abstract segment-to-prosody rules. These findings highlight a critical gap in the systematic modeling of fine-grained prosody within current TTS architectures and offer a novel perspective for assessing the naturalness of synthetic speech.

Technology Category

Application Category

📝 Abstract
This study proposes a segmental-level prosodic probing framework to evaluate neural TTS models' ability to reproduce consonant-induced f0 perturbation, a fine-grained segmental-prosodic effect that reflects local articulatory mechanisms. We compare synthetic and natural speech realizations for thousands of words, stratified by lexical frequency, using Tacotron 2 and FastSpeech 2 trained on the same speech corpus (LJ Speech). These controlled analyses are then complemented by a large-scale evaluation spanning multiple advanced TTS systems. Results show accurate reproduction for high-frequency words but poor generalization to low-frequency items, suggesting that the examined TTS architectures rely more on lexical-level memorization than on abstract segmental-prosodic encoding. This finding highlights a limitation in such TTS systems' ability to generalize prosodic detail beyond seen data. The proposed probe offers a linguistically informed diagnostic framework that may inform future TTS evaluation methods, and has implications for interpretability and authenticity assessment in synthetic speech.
Problem

Research questions and friction points this paper is trying to address.

neural TTS
F0 perturbation
prosodic generalization
segmental prosody
consonant-induced effects
Innovation

Methods, ideas, or system contributions that make the work stand out.

prosodic probing
consonant-induced F0 perturbation
neural TTS evaluation
segmental-prosodic modeling
lexical frequency stratification
🔎 Similar Papers
No similar papers found.
T
Tianle Yang
University at Buffalo, Department of Linguistics, Buffalo, 14260, NY, United States
C
Chengzhe Sun
University at Buffalo, Department of Computer Science and Engineering, Buffalo, 14260, NY, United States
P
Phil Rose
Australian National University, Emeritus Faculty, Canberra, 0200, ACT, Australia
C
Cassandra L. Jacobs
University at Buffalo, Department of Linguistics, Buffalo, 14260, NY, United States
S
Siwei Lyu
University at Buffalo, Department of Computer Science and Engineering, Buffalo, 14260, NY, United States