Towards a Phonology-Informed Evaluation of Multilingual TTS

๐Ÿ“… 2026-07-02
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
While modern text-to-speech (TTS) systems produce natural-sounding speech, they often fail to faithfully preserve phonemic contrasts that are critical for lexical and grammatical distinctionsโ€”a shortcoming undetected by conventional evaluation metrics such as Mean Opinion Score (MOS). This work proposes the first TTS evaluation framework integrating phonological knowledge, employing a phoneme classifier trained on human speech alongside phonological feature annotations (e.g., [+ATR]), acoustic cue analysis, and cross-domain transfer techniques to conduct fine-grained audits of synthetic speech. Applied to Assamese ATR vowel harmony, the framework reveals that approximately one-third of intended [+ATR] mid vowels are erroneously realized as [-ATR]. Notably, classification accuracy based on predicted phonological labels exceeds that of transcription-based labels, uncovering a systematic phonological gap between intended linguistic targets and actual TTS output.
๐Ÿ“ Abstract
Neural TTS systems can sound natural across languages, but naturalness does not guarantee the preservation of sound contrasts that distinguish words from their grammatical forms. Standard metrics like MOS do not test for this. We propose a classifier-based framework that audits TTS output against language-specific phonological patterns using human speech as a benchmark. Testing Assamese advanced tongue root (ATR) vowel harmony with Meta's MMS TTS, we show that a classifier trained on human speech transfers to synthesized speech with minimal loss. The faithfulness audit reveals that [+ATR] mid vowels are realized as [-ATR] in 1/3 tokens despite an underlying [+ATR] specification, a bias absent in human speech. At the word level, predicted ATR labels classify harmony more accurately than transcription labels, indicating a gap between intended and produced phonology. The framework offers task-specific diagnostics and generalizes to other phonological contrasts with measurable acoustic cues.
Problem

Research questions and friction points this paper is trying to address.

multilingual TTS
phonological contrast
faithfulness evaluation
vowel harmony
speech synthesis
Innovation

Methods, ideas, or system contributions that make the work stand out.

phonology-informed evaluation
classifier-based audit
ATR vowel harmony
multilingual TTS
phonological faithfulness
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
S
Sneha Ray Barman
Centre for Linguistic Science & Technology, IIT Guwahati
Neeraj Kumar Sharma
Neeraj Kumar Sharma
Ram Lal Anand College, University of Delhi
Explainable AIDigital WatermarkingMachine LearningTrust and Reputation SystemsFuzzy Systems
S
Shakuntala Mahanta
Centre for Linguistic Science & Technology, IIT Guwahati; Department of Humanities & Social Sciences, IIT Guwahati