The \`{I}r\`{o}y\`{i}nSpeech Text Corpus: 24,905 Curated Yor\`ub\'a Sentences for Speech and Language Technology

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the scarcity of Yoruba speech-text data, its predominant religious domain bias, and Unicode normalization defects affecting over 60% of existing texts. Through manual proofreading, Unicode standardization, adaptation of openly licensed materials, and neutral-register editing, this project overcomes current corpus limitations and achieves localization of non-proper nouns. The resulting dataset comprises 24,905 manually verified, tone-marked Yoruba sentences spanning news and original content, effectively rectifying encoding errors and optimizing reading clarity. This high-quality, validated prompt set directly supports downstream tasks including tone restoration, grapheme-to-phoneme (G2P) conversion, and frontend development for text-to-speech (TTS) synthesis.
📝 Abstract
\`{I}r\`{o}y\`{i}nSpeech is a 42-hour, 80-speaker Yor\`ub\'a read-speech corpus whose audio has been distributed by ELRA since 2024. This paper describes the release of its text component: 24,905 unique, hand-verified, tone-marked Yor\`ub\'a sentences (275,897 tokens; 15,687 types), curated in 2022 as recording prompts. Roughly 11,000 sentences were adapted from openly licensed news material; the remainder were written in-house to broaden coverage beyond the religious translation that dominates existing Yor\`ub\'a corpora. Every sentence was checked by hand for tone-mark accuracy, edited for read-aloud clarity and a neutral register, and localised so that non-Yor\`ub\'a personal and place names appear in Yor\`ub\'a form. Preparing the text for release surfaced systematic Unicode normalisation failures affecting more than 60% of lines (with precomposed and decomposed forms of the same letter co-occurring within single sentences) which we document and correct. The corpus supports diacritic restoration, grapheme-to-phoneme conversion, TTS front-end development and orthographic research, and serves as a validated prompt set for new recording.
Problem

Research questions and friction points this paper is trying to address.

Yorùbá
text corpus
tone marking
Unicode normalization
speech and language technology
Innovation

Methods, ideas, or system contributions that make the work stand out.

Yorùbá corpus
tone-marked text
Unicode normalization
diacritic restoration
low-resource language