🤖 AI Summary
It remains unclear whether current neural text-to-speech (TTS) systems can accurately model fine-grained segmental prosodic phenomena—specifically, consonant-induced fundamental frequency (F0) perturbations—and how well they generalize to low-frequency words. This work proposes a linguistically motivated, segment-level prosody probing framework, training Tacotron 2 and FastSpeech 2 models on the LJ Speech corpus and evaluating their F0 modeling capabilities through large-scale multi-system comparisons and lexical frequency-stratified analyses. Results indicate that while both systems perform well on high-frequency words, they exhibit limited generalization on low-frequency items, suggesting reliance on lexical memorization rather than abstract segment-to-prosody rules. These findings highlight a critical gap in the systematic modeling of fine-grained prosody within current TTS architectures and offer a novel perspective for assessing the naturalness of synthetic speech.
📝 Abstract
This study proposes a segmental-level prosodic probing framework to evaluate neural TTS models' ability to reproduce consonant-induced f0 perturbation, a fine-grained segmental-prosodic effect that reflects local articulatory mechanisms. We compare synthetic and natural speech realizations for thousands of words, stratified by lexical frequency, using Tacotron 2 and FastSpeech 2 trained on the same speech corpus (LJ Speech). These controlled analyses are then complemented by a large-scale evaluation spanning multiple advanced TTS systems. Results show accurate reproduction for high-frequency words but poor generalization to low-frequency items, suggesting that the examined TTS architectures rely more on lexical-level memorization than on abstract segmental-prosodic encoding. This finding highlights a limitation in such TTS systems' ability to generalize prosodic detail beyond seen data. The proposed probe offers a linguistically informed diagnostic framework that may inform future TTS evaluation methods, and has implications for interpretability and authenticity assessment in synthetic speech.