🤖 AI Summary
This work addresses the lack of human effort estimation in constructing speech datasets for low-resource, highly colloquial, and low-literacy languages—exemplified by Bambara (Mali). It presents the first systematic quantification of manual correction time required for ASR transcriptions: 30 hours per hour of audio in lab settings and 36 hours per hour of audio in field settings. Using ethnographic fieldwork and structured time-logging, the study engages native speakers to correct ASR outputs while rigorously isolating environmental variables. Key contributions are: (1) establishing the first human-effort benchmark for speech annotation in low-literacy languages; (2) empirically demonstrating that field conditions significantly increase correction complexity; and (3) providing a reusable cost-modelling framework and empirical evidence to guide NLP resource development for similar languages.
📝 Abstract
Creating speech datasets for low-resource languages is a critical yet poorly understood challenge, particularly regarding the actual cost in human labor. This paper investigates the time and complexity required to produce high-quality annotated speech data for a subset of low-resource languages, low literacy Predominately Oral Languages, focusing on Bambara, a Manding language of Mali. Through a one-month field study involving ten transcribers with native proficiency, we analyze the correction of ASR-generated transcriptions of 53 hours of Bambara voice data. We report that it takes, on average, 30 hours of human labor to accurately transcribe one hour of speech data under laboratory conditions and 36 hours under field conditions. The study provides a baseline and practical insights for a large class of languages with comparable profiles undertaking the creation of NLP resources.