SENS-ASR: Semantic Embedding injection in Neural-transducer for Streaming Automatic Speech Recognition

📅 2026-02-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work proposes a semantics-enhanced streaming automatic speech recognition (ASR) approach to address performance degradation caused by the absence of future contextual information in low-latency scenarios. By integrating a knowledge distillation–based sentence-embedding language model with a context-aware semantic module, the method extracts high-level semantic cues from historical frames and injects them into a neural Transformer architecture. This integration effectively mitigates the context deficiency inherent in chunk-based streaming processing. Experimental results on standard benchmarks demonstrate a significant reduction in word error rate, validating the efficacy and novelty of leveraging semantic guidance to enhance streaming transcription quality.

Technology Category

Natural Language Processing: Sentence-level Semantics, Textual Inference, etc.Machine Learning: Time-Series/Data StreamsData Mining & Knowledge Management: Data Stream Mining

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Personalized, context-aware and across-device searchGraph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphs
📝 Abstract
Many Automatic Speech Recognition (ASR) applications require streaming processing of the audio data. In streaming mode, ASR systems need to start transcribing the input stream before it is complete, i.e., the systems have to process a stream of inputs with a limited (or no) future context. Compared to offline mode, this reduction of the future context degrades the performance of Streaming-ASR systems, especially while working with low-latency constraint. In this work, we present SENS-ASR, an approach to enhance the transcription quality of Streaming-ASR by reinforcing the acoustic information with semantic information. This semantic information is extracted from the available past frame-embeddings by a context module. This module is trained using knowledge distillation from a sentence embedding Language Model fine-tuned on the training dataset transcriptions. Experiments on standard datasets show that SENS-ASR significantly improves the Word Error Rate on small-chunk streaming scenarios.
Problem

Research questions and friction points this paper is trying to address.

Streaming Automatic Speech Recognition
low-latency
future context
Word Error Rate
semantic information
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic Embedding
Streaming ASR
Neural Transducer
Knowledge Distillation
Context Module
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Youness Dkhissi
Orange Innovation, 4 Rue du Clos Courtel, 35510 Cesson-Sévigné, France
Valentin Vielzeuf
Valentin Vielzeuf
Orange Labs
E
Elys Allesiardo
Orange Innovation, 4 Rue du Clos Courtel, 35510 Cesson-Sévigné, France
Anthony Larcher
Anthony Larcher
Professor Le Mans Université
Speech processingSpeaker verifcationLanguage Identification