Cost Analysis of Human-corrected Transcription for Predominately Oral Languages

📅 2025-10-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of human effort estimation in constructing speech datasets for low-resource, highly colloquial, and low-literacy languages—exemplified by Bambara (Mali). It presents the first systematic quantification of manual correction time required for ASR transcriptions: 30 hours per hour of audio in lab settings and 36 hours per hour of audio in field settings. Using ethnographic fieldwork and structured time-logging, the study engages native speakers to correct ASR outputs while rigorously isolating environmental variables. Key contributions are: (1) establishing the first human-effort benchmark for speech annotation in low-literacy languages; (2) empirically demonstrating that field conditions significantly increase correction complexity; and (3) providing a reusable cost-modelling framework and empirical evidence to guide NLP resource development for similar languages.

Technology Category

Natural Language Processing: SpeechApplication Domains: Humanities & Computational Social ScienceHumans and AI: Crowd Sourcing and Human Computation

Application Category

Economics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web data
📝 Abstract
Creating speech datasets for low-resource languages is a critical yet poorly understood challenge, particularly regarding the actual cost in human labor. This paper investigates the time and complexity required to produce high-quality annotated speech data for a subset of low-resource languages, low literacy Predominately Oral Languages, focusing on Bambara, a Manding language of Mali. Through a one-month field study involving ten transcribers with native proficiency, we analyze the correction of ASR-generated transcriptions of 53 hours of Bambara voice data. We report that it takes, on average, 30 hours of human labor to accurately transcribe one hour of speech data under laboratory conditions and 36 hours under field conditions. The study provides a baseline and practical insights for a large class of languages with comparable profiles undertaking the creation of NLP resources.
Problem

Research questions and friction points this paper is trying to address.

Analyze human labor costs for transcribing oral languages
Evaluate ASR correction complexity for low-resource languages
Establish transcription baselines for low-literacy speech datasets
Innovation

Methods, ideas, or system contributions that make the work stand out.

Human correction of ASR transcriptions for oral languages
Field study with native transcribers analyzing 53 hours
Baseline cost of 30-36 hours labor per speech hour
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yacouba Diarra
RobotsMali AI4D Lab — @robotsmali.org
N
Nouhoum Souleymane Coulibaly
RobotsMali AI4D Lab — @robotsmali.org
Michael Leventhal
Michael Leventhal
RobotsMali
machine learninglow-resource languagesparallel computing architecturesautomata processingXML