Score
Designs, builds, or evaluates systems and workflows that convert spoken audio into textual outputs — orthographic or phonetic — including manual annotation and automated models for real-time, noise-robust, and high‑accuracy transcription. Work includes measuring and improving transcription accuracy, developing multilingual and cross‑lingual modeling, embeddings and transfer methods, integrating transcription with translation and retrieval pipelines, and managing recording/annotation and cross‑language evaluation processes.
This study investigates the trade-off between efficiency and accuracy in semi-automatic transcription for spoken language corpus construction. Through a two-stage experiment, it compares the performance of expert and novice transcribers on three types of Italian conversational data under both manual and ASR-assisted conditions. The work proposes an integrated analytical framework combining word-level alignment, quality evaluation metrics, and statistical modeling to systematically quantify behavioral differences across transcription workflows. Results demonstrate that ASR substantially increases transcription speed, yet its impact on accuracy varies depending on dialogue type, transcriber expertise, and workflow configuration. The findings provide empirical support for the development of the KIParla corpus, showing that a fine-tuned and optimized semi-automatic pipeline can effectively accelerate annotation while maintaining high transcription quality.
Speech recognition exhibits insufficient robustness across diverse audio formats (WAV/MP3/FLAC/OGG) and challenging acoustic conditions—including speaker accents, background noise, and domain-specific terminology. To address this, we propose an offline speech recognition system built upon the Vosk Toolkit, introducing—within this framework for the first time—a systematic, plug-and-play integration of customizable language models to enable zero-connectivity, low-resource domain adaptation. The system combines FFmpeg-based audio preprocessing with KaldiRecognizer-based decoding and outputs structured documents via python-docx. Evaluated on technical documentation and meeting recordings, it achieves an average 32% reduction in word error rate (WER) compared to generic models, demonstrating substantial accuracy gains. Moreover, it supports both real-time streaming and batch offline transcription, balancing practical applicability with broad generalizability.
Automatic speech recognition (ASR) for endangered languages in linguistic fieldwork suffers from extremely low transcription efficiency due to spontaneous, noisy recordings and critically scarce labeled data—often less than one hour. Method: This study systematically evaluates and optimizes multilingual pretrained models (MMS and XLS-R) across five typologically diverse endangered languages. We introduce linguistics-informed data cleaning, controllable few-shot benchmark construction, and a reproducible fine-tuning protocol. Contribution/Results: We establish, for the first time, clear applicability boundaries: MMS significantly outperforms XLS-R with <1 hour of data, while XLS-R only matches MMS performance beyond 1 hour. Our optimized pipelines yield production-ready ASR systems for all five languages, achieving several-fold improvements in transcription efficiency and directly alleviating the transcription bottleneck in language documentation.
This study addresses the limitations of traditional automatic speech recognition (ASR) evaluation, which relies heavily on word error rate (WER) and fails to capture the grammatical and semantic characteristics of transcription errors. To overcome this, the authors propose two novel metrics: Part-of-Speech Error Rate (POSER) and Embedding Error Rate (EmbER), which quantify ASR output quality from the perspectives of grammatical correctness and semantic fidelity, respectively. By integrating language model rescoring, part-of-speech tagging, and semantic distance computation based on word embeddings, they construct a multidimensional qualitative evaluation framework. Experimental results demonstrate that these new metrics effectively reveal the contribution of language models to improving linguistic quality in transcriptions, thereby compensating for WER’s insufficiency in linguistic analysis.
This work addresses the lack of human effort estimation in constructing speech datasets for low-resource, highly colloquial, and low-literacy languages—exemplified by Bambara (Mali). It presents the first systematic quantification of manual correction time required for ASR transcriptions: 30 hours per hour of audio in lab settings and 36 hours per hour of audio in field settings. Using ethnographic fieldwork and structured time-logging, the study engages native speakers to correct ASR outputs while rigorously isolating environmental variables. Key contributions are: (1) establishing the first human-effort benchmark for speech annotation in low-literacy languages; (2) empirically demonstrating that field conditions significantly increase correction complexity; and (3) providing a reusable cost-modelling framework and empirical evidence to guide NLP resource development for similar languages.
This study addresses the challenge faced by field linguists who, due to limited technical expertise, struggle to leverage multilingual automatic speech recognition (ASR) systems, making audio transcription a bottleneck in language documentation. To overcome this, the work proposes the first no-code ASR fine-tuning workflow tailored for linguists, seamlessly integrating the ELAN annotation environment with the multilingual Whisper model via a cloud-based platform to enable iterative refinement. A novel cold-start transcription prioritization strategy is introduced, which selects utterances based on linguistic richness—such as lexical diversity and phonemic coverage—rather than acoustic quality. Experiments on three low-resource languages from Vanuatu demonstrate that this approach substantially improves both transcription efficiency and accuracy, rapidly reducing character error rate (CER) even under noisy recording conditions.
This study addresses the challenge of achieving robust phonetic transcription in low-resource scenarios involving non-standard dialects and atypical speech—such as non-native and post-stroke dysarthric utterances—where expert phonetic annotation is prohibitively expensive. The authors systematically investigate the interplay between grapheme-to-phoneme (G2P) models and human supervision, revealing on an 80-hour multi-variety speech benchmark that G2P augmentation benefits performance only when fewer than 20–30 hours of manual transcriptions are available; beyond this threshold, it degrades cross-dialect generalization. To overcome this limitation, they propose replacing G2P with ASR-based pretraining and introduce a weighted phonetic feature error rate for evaluation. This approach yields substantial gains on both non-native and aphasic speech, reducing error rates by a factor of 2.3 compared to prior systems.
This work addresses the challenge of conducting fine-grained, part-of-speech (PoS)-level error analysis in automatic speech recognition (ASR) for non-Latin script languages, where reliable word-level alignment is often unattainable. To overcome this limitation, the authors propose a language-agnostic automatic alignment framework that, for the first time, uniformly supports the three major writing systems—Abugida, Alphabetic, and Abjad. By integrating a general-purpose sequence alignment algorithm with standard PoS taggers, the framework enables a scalable and reproducible pipeline for PoS-level ASR error analysis. This approach effectively removes linguistic barriers in ASR diagnostics for non-Latin scripts and demonstrates practical utility: in multilingual experiments, insights derived from the analysis were successfully fed back into ASR training, yielding significant reductions in word error rate (WER).
Traditional automatic speech recognition evaluation metrics, such as word error rate (WER) and character error rate (CER), fail to capture human perception of errors and neglect linguistic and semantic influences. This work proposes a novel paradigm that embeds any perception-oriented evaluation metric into the minimum edit distance (minED) framework to produce an intuitively interpretable equivalent error rate. For the first time, this approach translates human perceptual modeling into a comprehensible error rate format, enabling quantification of error severity from the perspective of human understanding. The resulting metric not only aligns closely with human judgments but also effectively identifies recognition errors that critically impact semantic comprehension.