Factorized Delayed Streams Modeling for LLM-based Streaming ASR

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of special token interference in text prediction and high computational overhead in large language model (LLM)-based streaming speech recognition. To this end, we propose the Factorized Delay Streaming Model (F-DSM). This method decouples the waiting mechanism from the large-vocabulary distribution via probabilistic factorization, eliminates redundant word-start tokens, and introduces a softmax skipping strategy to reduce unnecessary computations. Experimental results demonstrate that F-DSM improves recognition accuracy while maintaining stable text perplexity. Furthermore, it significantly reduces GPU memory consumption and accelerates inference. Overall, this work provides an efficient delay modeling solution for LLM-based streaming automatic speech recognition.
📝 Abstract
Delayed Streams Modeling (DSM) enables LLM-based streaming automatic speech recognition (ASR) by aligning acoustic and text streams on a common timeline. DSM adds the padding tokenand the word-start tokento the LLM vocabulary and predicts them together with normal text tokens using the same softmax. We first show thatcan be removed while maintaining competitive recognition performance. Based on this result, we propose Factorized DSM (F-DSM), which separates the waiting probability forfrom the distribution over the original LLM vocabulary. This factorization removes ASR-specific tokens from the text prediction space and allows the large-vocabulary softmax to be skipped on waiting steps. Experiments on the Corpus of Spontaneous Japanese and LibriSpeech show that F-DSM achieves better recognition performance than DSM. It also greatly reduces GPU memory use while maintaining similar training throughput, provides a small inference speed improvement through softmax skipping, and reduces the degradation in text-only perplexity observed with DSM.
Problem

Research questions and friction points this paper is trying to address.

Streaming ASR
Large Language Model
Delayed Streams Modeling
Softmax factorization
Memory efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Factorized Delayed Streams Modeling
Streaming ASR
Large Language Model
Softmax Skipping
Memory Efficiency
🔎 Similar Papers