Can Prosodic Style Be Inferred from Text Alone? Evidence from Unsupervised Acoustic Clusters

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses whether text alone suffices to infer speech prosodic style, challenging a default assumption in text-to-speech (TTS) synthesis by examining the explanatory power of text over prosody. To this end, we propose a falsifiable hypothesis-testing framework that introduces majority-class and bag-of-words baselines to disentangle sentence-length confounds. Through large-scale controlled corpus analyses combining five acoustic clustering methods, twelve text embedding types, and tree-based probing techniques, we reveal that textual prediction of prosody relies primarily on lexical choice rather than deep semantics, with sentence-level embeddings yielding only marginal gains. These findings demonstrate that reference-free style selection should not depend excessively on textual signals, offering new perspectives for understanding prosody generation mechanisms in TTS systems.
📝 Abstract
Much of expressive text-to-speech research rests on an untested assumption that written text carries enough information to select an appropriate prosodic style for its delivery. Text-predicted style models improve listener preference, and expressive-appropriateness evaluation presupposes that context constrains style, yet neither measures the assumption itself. This paper tests it as a falsifiable hypothesis against style labels derived from acoustics alone. For each of six speakers in a 1,200-hour conversational corpus, utterances are clustered in the spaces of five speech models, including a prosody-only control, and the cluster of held-out utterances is predicted from twelve text embedding models. Three controls are applied: utterance length is erased from the speech embeddings; accuracy is scored against the majority-class floor of unbalanced clusters rather than uniform chance; and a bag-of-words baseline measures word identity alone. Text predicts the cluster above that floor for all six speakers (+0.111 top-3 accuracy), but bag-of-words achieves three quarters of this. Sentence embeddings add only +0.026, largest for encoders not trained for sentence semantics and reversed by tree-based probes for all others. Acoustic clusters are not compact in text embedding space in any of 360 configurations. The prosody-only space weakens the association for five speakers, but not for the speaker showing it most strongly. Text thus informs these delivery clusters mainly through word choice, whether as a cue to prosody or as a marker of topic and recording situation, and reference-free style selection cannot assume more.
Problem

Research questions and friction points this paper is trying to address.

prosodic style
text-to-speech
expressive speech
acoustic clustering
text embeddings
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unsupervised acoustic clustering
Prosodic style prediction
Text embeddings
Expressive text-to-speech
Falsifiable hypothesis testing
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Abdul Rehman
Abdul Rehman
Centre for Animal Disease Modeling and Surveillance, Vet School, University of California Davis
Veterinary epidemiologyOne HealthZoonotic pathogensAntimicrobial resistanceVector-borne
J
Jian-Jun Zhang
National Centre for Computer Animation, Bournemouth University, U.K.
X
Xiaosong Yang
National Centre for Computer Animation, Bournemouth University, U.K.