Evaluation of real-time transcriptions using end-to-end ASR models

📅 2024-09-09
🏛️ arXiv.org
📈 Citations: 6
✨ Influential: 0
📄 PDF
🤖 AI Summary
ASR systems face a fundamental trade-off between low latency and high accuracy in real-time interpretation scenarios. This paper investigates the impact of audio segmentation strategies on the latency and performance of end-to-end ASR models, proposing a feedback-driven dynamic segmentation algorithm. Unlike fixed-interval segmentation—which minimizes latency but substantially degrades WER—or VAD-based segmentation—which optimizes WER at the cost of maximal latency—our approach dynamically adjusts segment boundaries using real-time recognition feedback to balance both objectives. Experimental results show that our method reduces end-to-end latency by 1.5–2 seconds relative to VAD-based segmentation while increasing WER by only 2–4 percentage points. The proposed feedback-based segmentation is empirically validated on real-time transcription tasks, demonstrating robustness and practicality. This work establishes a deployable paradigm for low-latency, high-fidelity end-to-end ASR systems.

Technology Category

Computer Vision: SegmentationMachine Learning: Hardware-aware MLNatural Language Processing: Interpretability, Analysis, and Evaluation of NLP Models

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsSystems and Infrastructure for Web, Mobile and WoT: Web performance, measurement, and characterization
📝 Abstract
Automatic Speech Recognition (ASR) or Speech-to-text (STT) has greatly evolved in the last few years. Traditional architectures based on pipelines have been replaced by joint end-to-end (E2E) architectures that simplify and streamline the model training process. In addition, new AI training methods, such as weak-supervised learning have reduced the need for high-quality audio datasets for model training. However, despite all these advancements, little to no research has been done on real-time transcription. In real-time scenarios, the audio is not pre-recorded, and the input audio must be fragmented to be processed by the ASR systems. To achieve real-time requirements, these fragments must be as short as possible to reduce latency. However, audio cannot be split at any point as dividing an utterance into two separate fragments will generate an incorrect transcription. Also, shorter fragments provide less context for the ASR model. For this reason, it is necessary to design and test different splitting algorithms to optimize the quality and delay of the resulting transcription. In this paper, three audio splitting algorithms are evaluated with different ASR models to determine their impact on both the quality of the transcription and the end-to-end delay. The algorithms are fragmentation at fixed intervals, voice activity detection (VAD), and fragmentation with feedback. The results are compared to the performance of the same model, without audio fragmentation, to determine the effects of this division. The results show that VAD fragmentation provides the best quality with the highest delay, whereas fragmentation at fixed intervals provides the lowest quality and the lowest delay. The newly proposed feedback algorithm exchanges a 2-4% increase in WER for a reduction of 1.5-2s delay, respectively, to the VAD splitting.
Problem

Research questions and friction points this paper is trying to address.

Measuring ASR system delay for real-time interpretation scenarios
Addressing latency mismatch between ASR transcription and human interpretation
Validating ASR usability in live settings requiring immediate translation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Proposes new ASR latency measurement method
Validates ASR usability for live interpretation
Measures speech-to-transcription delivery delay
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Carlos Arriaga
A
Alejandro Pozo
J
J. Conde
Á
Álvaro Alonso