Yo-ByT5: Efficient and High-Fidelity Diacritic Restoration for Yorùbá

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the severe lexical ambiguity in Yoruba caused by missing diacritics, which significantly constrains its natural language processing applications. To mitigate this issue, we propose a byte-level automatic diacritic restoration model fine-tuned on ByT5-small. Leveraging byte-level pretraining and a Transformer architecture, the model directly processes character sequences without requiring tokenization. Notably, it achieves performance comparable to mT5-base while utilizing only half the parameters, thereby substantially reducing computational overhead. Experimental results demonstrate that the proposed model attains a diacritic error rate (DER) of 10.14% and a character error rate (CER) of 3.48%. These findings indicate that our approach effectively balances high text fidelity with computational efficiency, making it particularly suitable for low-resource scenarios.
📝 Abstract
Yorùbá is a widely spoken tonal language that depends on diacritics to avoid lexical ambiguity. However, it is often written without these diacritics, thereby hindering downstream Natural Language Processing (NLP) tasks. In this paper, we introduce Yo-ByT5, a byte-level Automatic Diacritic Restoration (ADR) model fine-tuned from ByT5-small. We evaluate Yo-ByT5 alongside five publicly released Yorùbá ADR models and one open-weight large language model (LLM) on the YAD benchmark under a consistent protocol. Our results demonstrate that Yo-ByT5 matches the performance of the strongest existing model, mT5-base, with a DER of 10.14% and a CER of 3.48%. Furthermore, it exhibits superior text fidelity despite using approximately half the parameter count of mT5-base. We also release our training code and model outputs, as well as call for the development of a larger, purpose-built benchmark for Yorùbá diacritic restoration.
Problem

Research questions and friction points this paper is trying to address.

Yorùbá
Diacritic Restoration
Natural Language Processing
Tonal Language
Innovation

Methods, ideas, or system contributions that make the work stand out.

Byte-level model
Diacritic restoration
ByT5
Parameter efficiency
Yorùbá
🔎 Similar Papers
No similar papers found.