Prosody-to-Text: Predicting text from low-pass filtered speech

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study presents the first systematic investigation of the overlooked inverse prediction problem of "prosody-to-text," exploring the feasibility of recovering original sentences solely from speech prosody. Methodologically, low-pass filtering is employed to extract the first twelve Mel-spectral components, preserving low-frequency prosodic features, which are then used to fine-tune a Whisper model for mapping prosody to lexical content. Experimental results demonstrate that the proposed approach reduces the word error rate to 36%, with 10% of sentences perfectly reconstructed. Furthermore, given a prefix context, next-word prediction accuracy reaches 79%. These findings reveal a strong correlation between low-frequency speech signals and lexical content, offering a novel paradigm for cross-modal speech-language understanding.
📝 Abstract
While predicting prosody from text is an established task in the field, the opposite direction, predicting text that fits a given prosodic pattern, remains largely overlooked. We find this unfortunate, because this opposite direction could lead to some very interesting use cases. Therefore, in this paper, we make the first steps in the prosody-to-text direction by inves- tigating how much of the original sentence can be recovered from its prosodic pattern. To this end, we fine-tune the Whis- per model using only the 12 lowest Mel bins (low-pass filter with approximately 450Hz cutoff), and obtain surprisingly accurate results (WER 36%), with 10% of utterances be- ing recovered perfectly, and 40% of utterances having Word Error Rate at or below 25%. We also find that, given the correct prefix, the next token was predicted correctly in 79% of cases. Our results suggest that the relationship between low-frequency speech features and lexical content is much stronger than previously thought, and we believe that direct- ing more attention to this topic might open the door to new applications, such as using prosody to guide text generation of modern LLMs
Problem

Research questions and friction points this paper is trying to address.

Prosody-to-Text
Low-pass Filtered Speech
Speech Recognition
Lexical Content Recovery
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prosody-to-Text
Whisper fine-tuning
Low-pass filtered speech
Speech representation
Text generation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
David Porteš
Natural Language Processing Centre, Faculty of Informatics, Masaryk University, Botanicka 68a, 602 00 Brno, Czech Republic
Aleš Horák
Aleš Horák
NLP Centre, Faculty of Informatics, Masaryk University
Natural Language ProcessingComputational LinguisticsInformation RetrievalArtificial Intelligence