Score
Implements procedures and tooling to compute Word Error Rate (WER) between reference and hypothesis text by performing tokenization/normalization, sequence alignment, and counting insertions, deletions, and substitutions. Builds evaluation reports and aggregated statistics to compare models or training conditions and supports language-specific tokenization and edge cases (e.g., punctuation, casing, empty hypotheses).
This work addresses the overestimation of word error rate (WER) in multi-script automatic speech recognition (ASR) scenarios, where reference and hypothesis transcripts employ different writing systems—such as romanized versus native scripts—leading to inflated error counts that conflate true recognition failures with script mismatches. To resolve this, the authors propose SN-WER, a training-free evaluation metric that incorporates language-specific script normalization to transcribe both reference and hypothesis texts into a standardized script prior to WER computation. Evaluated across five Indian languages, two datasets, and three ASR models, SN-WER reduces apparent performance gaps between systems by up to 12%, mitigates artificial error inflation from manual romanization by 67% in stress tests, and maintains a word collision rate below 0.1%, thereby significantly enhancing the fairness and robustness of ASR evaluation in multi-script settings.
Traditional automatic speech recognition evaluation metrics, such as word error rate (WER) and character error rate (CER), fail to capture human perception of errors and neglect linguistic and semantic influences. This work proposes a novel paradigm that embeds any perception-oriented evaluation metric into the minimum edit distance (minED) framework to produce an intuitively interpretable equivalent error rate. For the first time, this approach translates human perceptual modeling into a comprehensible error rate format, enabling quantification of error severity from the perspective of human understanding. The resulting metric not only aligns closely with human judgments but also effectively identifies recognition errors that critically impact semantic comprehension.
This study addresses the limitations of traditional automatic speech recognition (ASR) evaluation, which relies heavily on word error rate (WER) and fails to capture the grammatical and semantic characteristics of transcription errors. To overcome this, the authors propose two novel metrics: Part-of-Speech Error Rate (POSER) and Embedding Error Rate (EmbER), which quantify ASR output quality from the perspectives of grammatical correctness and semantic fidelity, respectively. By integrating language model rescoring, part-of-speech tagging, and semantic distance computation based on word embeddings, they construct a multidimensional qualitative evaluation framework. Experimental results demonstrate that these new metrics effectively reveal the contribution of language models to improving linguistic quality in transcriptions, thereby compensating for WER’s insufficiency in linguistic analysis.
Conventional automatic speech recognition (ASR) evaluation using word error rate (WER) fails to capture the practical impact of ASR errors on downstream large language model (LLM)-driven tasks. Method: We propose a task-oriented ASR evaluation framework that (1) systematically classifies ASR error types and analyzes their contextual reparability within LLM prompts; (2) defines a multidimensional metric integrating semantic severity of errors, LLM-based correction success rate, and end-task completion accuracy; and (3) validates the framework empirically on representative speech-to-LLM pipelines—including voice command execution and meeting summary generation. Results: Our framework significantly outperforms WER in reflecting ASR effectiveness in real-world LLM applications. It provides an interpretable, quantifiable assessment grounded in downstream task performance, enabling principled, task-aware ASR model development and optimization.
To address the low efficiency and insufficient accuracy of Word Error Rate (WER) estimation for automatic speech recognition (ASR) systems in label-free evaluation scenarios, this paper proposes an efficient unsupervised WER estimation algorithm. The method fuses self-supervised speech representations (wav2vec 2.0) and text representations (BERT) via mean pooling, then employs a lightweight regression model to directly predict WER—bypassing computationally expensive alignment procedures. Evaluated on the Ted-Lium3 benchmark, the approach achieves a 14.10% reduction in root-mean-square error (RMSE) and a 1.22 percentage-point improvement in Pearson correlation coefficient over prior methods, while attaining a real-time factor of 3.4×. This demonstrates a favorable trade-off between accuracy and inference speed. The proposed framework provides a scalable, fully unsupervised solution for large-scale ASR evaluation without ground-truth transcriptions.
This work addresses the limitations of conventional automatic speech recognition (ASR) evaluation metrics—such as word error rate (WER)—which exhibit poor performance in low-resource languages and hinder fair cross-lingual comparisons. To overcome these challenges, the authors propose an enhanced WER evaluation framework that integrates language-specific text normalization, compound word detection, and a fine-grained token-level Levenshtein distance-based alignment mechanism. Notably, this is the first open-source implementation to incorporate an alignment representation capable of embedding metadata. Evaluated across 52 languages, the proposed method achieves absolute WER reductions of up to 25%, substantially improving the robustness and fairness of cross-lingual ASR evaluation.