🤖 AI Summary
This work addresses the scarcity of Yoruba speech-text data, its predominant religious domain bias, and Unicode normalization defects affecting over 60% of existing texts. Through manual proofreading, Unicode standardization, adaptation of openly licensed materials, and neutral-register editing, this project overcomes current corpus limitations and achieves localization of non-proper nouns. The resulting dataset comprises 24,905 manually verified, tone-marked Yoruba sentences spanning news and original content, effectively rectifying encoding errors and optimizing reading clarity. This high-quality, validated prompt set directly supports downstream tasks including tone restoration, grapheme-to-phoneme (G2P) conversion, and frontend development for text-to-speech (TTS) synthesis.
📝 Abstract
\`{I}r\`{o}y\`{i}nSpeech is a 42-hour, 80-speaker Yor\`ub\'a read-speech corpus whose audio has been distributed by ELRA since 2024. This paper describes the release of its text component: 24,905 unique, hand-verified, tone-marked Yor\`ub\'a sentences (275,897 tokens; 15,687 types), curated in 2022 as recording prompts. Roughly 11,000 sentences were adapted from openly licensed news material; the remainder were written in-house to broaden coverage beyond the religious translation that dominates existing Yor\`ub\'a corpora. Every sentence was checked by hand for tone-mark accuracy, edited for read-aloud clarity and a neutral register, and localised so that non-Yor\`ub\'a personal and place names appear in Yor\`ub\'a form. Preparing the text for release surfaced systematic Unicode normalisation failures affecting more than 60% of lines (with precomposed and decomposed forms of the same letter co-occurring within single sentences) which we document and correct. The corpus supports diacritic restoration, grapheme-to-phoneme conversion, TTS front-end development and orthographic research, and serves as a validated prompt set for new recording.