A Sociolinguistic Analysis of Automatic Speech Recognition Bias in Newcastle English

📅 2026-03-25
📈 Citations: 0
Influential: 0
📄 PDF

career value

175K/year
🤖 AI Summary
This study addresses the systemic social biases in commercial automatic speech recognition (ASR) systems when processing non-mainstream dialects such as Newcastle English. For the first time, a sociolinguistic perspective is systematically integrated into ASR bias analysis, examining over 3,000 spontaneous speech transcription errors from the DECTE corpus through a linguistic classification framework and acoustic analysis. The findings reveal that phonological variation—including vowel quality, glottalization, and local lexical items—is a primary source of misrecognition, with significantly higher error rates observed for male speakers and those at the age extremes. These results demonstrate that ASR errors follow structured sociodemographic patterns, underscoring the critical importance of incorporating dialectal variation and community-specific speech data to develop more equitable voice technologies.

Technology Category

Application Category

📝 Abstract
Automatic Speech Recognition (ASR) systems are widely used in everyday communication, education, healthcare, and industry, yet their performance remains uneven across speakers, particularly when dialectal variation diverges from the mainstream accents represented in training data. This study investigates ASR bias through a sociolinguistic analysis of Newcastle English, a regional variety of North-East England that has been shown to challenge current speech recognition technologies. Using spontaneous speech from the Diachronic Electronic Corpus of Tyneside English (DECTE), we evaluate the output of a state-of-the-art commercial ASR system and conduct a fine-grained analysis of more than 3,000 transcription errors. Errors are classified by linguistic domain and examined in relation to social variables including gender, age, and socioeconomic status. In addition, an acoustic case study of selected vowel features demonstrates how gradient phonetic variation contributes directly to misrecognition. The results show that phonological variation accounts for the majority of errors, with recurrent failures linked to dialect-specific features like vowel quality and glottalisation, as well as local vocabulary and non-standard grammatical forms. Error rates also vary across social groups, with higher error frequencies observed for men and for speakers at the extremes of the age spectrum. These findings indicate that ASR errors are not random but socially patterned and can be explained from a sociolinguistic perspective. Thus, the study demonstrates the importance of incorporating sociolinguistic expertise into the evaluation and development of speech technologies and argues that more equitable ASR systems require explicit attention to dialectal variation and community-based speech data.
Problem

Research questions and friction points this paper is trying to address.

Automatic Speech Recognition
dialectal variation
sociolinguistic bias
Newcastle English
speech recognition errors
Innovation

Methods, ideas, or system contributions that make the work stand out.

sociolinguistic analysis
dialectal variation
ASR bias
phonetic gradient
socially patterned errors