Automatic Speech Recognition for Low-Resource Sinhala: A Critical Review of Methods, Challenges, and Future Directions

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of data scarcity, morphological complexity, and the absence of standardized benchmarks in Sinhala automatic speech recognition (ASR) by presenting the first comprehensive critical review tracing the field’s evolution from hidden Markov models to self-supervised architectures. Methodologically, it systematically compares multilingual low-resource strategies and rigorously evaluates pretrained models—including wav2vec 2.0, XLS-R, and Whisper—alongside transfer learning frameworks. The primary contributions include identifying six distinct research gaps and proposing a systematic improvement framework tailored for morphologically rich languages. Furthermore, this work exposes the incomparability of existing word error rate (WER) metrics and quantitatively demonstrates that corpus correction yields an 18.1% relative error reduction, thereby establishing a standardized direction and roadmap for future research in this domain.
📝 Abstract
Automatic speech recognition (ASR) for low-resource languages remains a major challenge. Sinhala, the primary language of Sri Lanka with about 16 million speakers, illustrates the difficulty: agglutinative morphology, a 54-phoneme inventory, subject-object-verb (SOV) syntax and scarce annotated speech data limit both conventional and modern ASR systems. This paper presents the first critical review of Sinhala ASR research, tracing its development from Hidden Markov Models (HMMs) through deep neural networks to self-supervised pre-trained models such as wav2vec 2.0, XLS-R, Whisper and Massively Multilingual Speech (MMS). We compare existing Sinhala systems with related low-resource ASR work on Tamil, Malayalam and Hindi in terms of architecture, training data, word error rate (WER) and robustness to real-world acoustic conditions, and we assess self-supervised and transfer learning as responses to scarce labeled data. We show that most reported WERs are not directly comparable because they differ in corpus, data split and scoring, and that the only controlled comparison in the literature attributes an 18.1% relative WER reduction to corpus correction alone. We also discuss context-aware ASR that draws on phonological, syntactic and semantic knowledge. We identify six research gaps: (1) the lack of large annotated corpora covering multiple dialects and acoustic conditions; (2) weak contextual modeling of Sinhala morphosyntax; (3) high WER in real-world conditions; (4) the absence of standardized benchmarks; (5) the lack of parameter-efficient fine-tuning studies; and (6) the absence of annotated code-switched Sinhala-English speech resources. We outline a research agenda to address these gaps, intended as a roadmap for researchers working on Sinhala and other morphologically rich languages.
Problem

Research questions and friction points this paper is trying to address.

Automatic Speech Recognition
Low-Resource Languages
Sinhala
Morphologically Rich Languages
Code-Switching
Innovation

Methods, ideas, or system contributions that make the work stand out.

Low-Resource ASR
Self-Supervised Learning
Sinhala Speech Recognition
Transfer Learning
Context-Aware ASR
🔎 Similar Papers
No similar papers found.