BaLEEN: Biasing with Latent Encoded Entities for Context-Aware ASR

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of transcribing proper nouns and rare entities in automatic speech recognition (ASR) by proposing the BaLEEN framework. This method leverages a pretrained language model with a Perceiver bottleneck architecture to encode contextual keywords, which are then mapped into latent vectors via a hypernetwork and injected into a frozen CTC-ASR encoder. This design enables dynamic context adaptation without fine-tuning. Functioning as a plug-and-play module, BaLEEN introduces no additional computational overhead during inference. Experimental results demonstrate that the proposed framework reduces the keyword miss rate by 8.7% while achieving relative improvements of 21% and 28% in word error rate and character error rate, respectively.
📝 Abstract
Transcribing domain-specific entities and rare proper nouns remains a major challenge in automatic speech recognition (ASR). In this paper, we propose BaLEEN (Biasing with Latent Encoded Entities), a lightweight, hypernetwork-based framework for dynamic contextual adaptation without fine-tuning the underlying ASR model. BaLEEN encodes variable-length contextual keywords using a pretrained language model, compresses them into a fixed sequence of latent vectors via a Perceiver bottleneck, and injects context-dependent bias vectors directly into the intermediate encoder representations of the ASR model. Because both the language model and the backbone ASR model remain entirely frozen during training, BaLEEN operates as a plug-and-play adapter that incurs zero computational overhead at inference time when context biases are precomputed. We evaluate our method on a CTC-based ASR model using a Wikipedia-derived corpus with annotated named entities and synthetic speech. Experimental results demonstrate that BaLEEN reduces keyword miss rate by 8.7% on the test set relative to the unbiased baseline while simultaneously improving overall word error rate by 21% and character error rate by 28%.
Problem

Research questions and friction points this paper is trying to address.

Automatic Speech Recognition
Domain-specific Entities
Rare Proper Nouns
Contextual Adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Context-Aware ASR
Hypernetwork
Perceiver Bottleneck
Plug-and-Play Adapter
Latent Encoded Entities
🔎 Similar Papers
No similar papers found.