🤖 AI Summary
This study addresses the systemic failures of automatic speech recognition (ASR) systems on low-resource, Indigenous, and non-standard language varieties, which reproduce colonial linguistic hierarchies. Integrating theories of linguistic capital, raciolinguistic ideologies, and decolonial computing, the work proposes a participatory framework that positions affected communities as co-designers, evaluators, and governance partners. It introduces an innovative “triple harm” taxonomy—comprising misrecognition, misalignment, and mistrust—and a seven-layer contextual model of linguistic diversity, embedding critical language policy perspectives directly into ASR design and evaluation. The resulting culturally competent ASR framework, accompanied by a minimal auditing protocol, significantly enhances the visibility, representation, and agency of marginalized language communities in speech technologies.
📝 Abstract
This paper focuses on automatic speech recognition (ASR) and ASR-mediated voice interfaces that shape access to public services, healthcare, and education. We argue that persistent failures for low-resource, Indigenous, and non-standard language varieties are not only technical errors, but also implicit linguistic policies that reproduce colonial language hierarchies. Drawing on linguistic capital, raciolinguistic ideology, language policy research, and decolonial computing, we show how data, metrics, and model priors determine whose voices become machine-legible. We introduce the Three Harms (3M) taxonomy---Misrecognition, Misalignment, and Mistrust---and a seven-layer situatedness model for linguistic diversity in ASR and ASR-mediated voice interfaces. We then propose a participatory framework and minimum audit protocol for culturally competent ASR, positioning affected communities as co-designers, evaluators, and governance partners.