🤖 AI Summary
High-quality benchmark data for automatic naturalness assessment of Spanish text-to-speech (TTS) systems is lacking. Method: This paper introduces SPANAT—the first publicly available, diverse, and large-scale Spanish TTS naturalness evaluation dataset—comprising 4,326 audio samples from 52 TTS systems and human speech, subjectively annotated via Mean Opinion Score (MOS) by 92 participants following ITU-T P.807. We further propose two novel modeling paradigms leveraging self-supervised speech representations: (i) fine-tuning English-pretrained models (e.g., wav2vec 2.0), and (ii) attaching lightweight downstream networks to frozen encoders for cross-lingual transfer. Contribution/Results: Our best model achieves a mean absolute error (MAE) of 0.80 on the five-point MOS scale, validating both the dataset’s utility and the efficacy of the proposed approaches. SPANAT establishes a reproducible benchmark and advances methodology for naturalness assessment in low-resource languages.
📝 Abstract
This work addresses the development of a database for the automatic assessment of text-to-speech (TTS) systems in Spanish, aiming to improve the accuracy of naturalness prediction models. The dataset consists of 4,326 audio samples from 52 different TTS systems and human voices and is, up to our knowledge, the first of its kind in Spanish. To label the audios, a subjective test was designed based on the ITU-T Rec. P.807 standard and completed by 92 participants. Furthermore, the utility of the collected dataset was validated by training automatic naturalness prediction systems. We explored two approaches: fine-tuning an existing model originally trained for English, and training small downstream networks on top of frozen self-supervised speech models. Our models achieve a mean absolute error of 0.8 on a five-point MOS scale. Further analysis demonstrates the quality and diversity of the developed dataset, and its potential to advance TTS research in Spanish.