A Dataset for Automatic Assessment of TTS Quality in Spanish

📅 2025-07-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
High-quality benchmark data for automatic naturalness assessment of Spanish text-to-speech (TTS) systems is lacking. Method: This paper introduces SPANAT—the first publicly available, diverse, and large-scale Spanish TTS naturalness evaluation dataset—comprising 4,326 audio samples from 52 TTS systems and human speech, subjectively annotated via Mean Opinion Score (MOS) by 92 participants following ITU-T P.807. We further propose two novel modeling paradigms leveraging self-supervised speech representations: (i) fine-tuning English-pretrained models (e.g., wav2vec 2.0), and (ii) attaching lightweight downstream networks to frozen encoders for cross-lingual transfer. Contribution/Results: Our best model achieves a mean absolute error (MAE) of 0.80 on the five-point MOS scale, validating both the dataset’s utility and the efficacy of the proposed approaches. SPANAT establishes a reproducible benchmark and advances methodology for naturalness assessment in low-resource languages.

Technology Category

Natural Language Processing: SpeechMachine Learning: Large Multimodal Models (LMMs)Cognitive Modeling & Cognitive Systems: Neural Spike Coding

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasets
📝 Abstract
This work addresses the development of a database for the automatic assessment of text-to-speech (TTS) systems in Spanish, aiming to improve the accuracy of naturalness prediction models. The dataset consists of 4,326 audio samples from 52 different TTS systems and human voices and is, up to our knowledge, the first of its kind in Spanish. To label the audios, a subjective test was designed based on the ITU-T Rec. P.807 standard and completed by 92 participants. Furthermore, the utility of the collected dataset was validated by training automatic naturalness prediction systems. We explored two approaches: fine-tuning an existing model originally trained for English, and training small downstream networks on top of frozen self-supervised speech models. Our models achieve a mean absolute error of 0.8 on a five-point MOS scale. Further analysis demonstrates the quality and diversity of the developed dataset, and its potential to advance TTS research in Spanish.
Problem

Research questions and friction points this paper is trying to address.

Develops a Spanish TTS quality assessment dataset
Validates dataset via naturalness prediction models
First large-scale Spanish TTS evaluation resource
Innovation

Methods, ideas, or system contributions that make the work stand out.

Developed Spanish TTS quality assessment dataset
Used ITU-T P.807 standard for labeling
Fine-tuned English model for Spanish prediction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Alejandro Sosa Welford
Ingeniería de Sonido, Universidad Nacional de Tres de Febrero, Argentina; Cognitive Neuroscience Center, Universidad de San Andrés, Argentina
L
Leonardo Pepino
Instituto de Investigacion en Ciencias de la Computación (ICC), CONICET-UBA, Argentina; Departamento de Computación, Universidad de Buenos Aires, Argentina