🤖 AI Summary
CTC-based ASR models offer efficient non-autoregressive decoding but suffer from weak language modeling capability, limiting recognition accuracy. To address this, we propose an LLM-driven intermediate loss regularization framework: leveraging the LLaMA embedding space as semantic supervision, we map intermediate representations from a Conformer encoder into the LLM’s semantic space via a learnable projection layer and jointly optimize with a causal language modeling loss. Crucially, this approach enhances speech features with linguistic knowledge without modifying the CTC decoding pipeline. Experiments on LibriSpeech, TEDLIUM2, and WSJ demonstrate substantial WER reductions, achieving state-of-the-art performance among CTC-based methods with negligible computational overhead. Our key contribution is the first use of an LLM’s semantic space as an intermediate supervision signal, enabling lightweight, efficient, and end-to-end language-aware regularization.
📝 Abstract
End-to-end (E2E) automatic speech recognition (ASR) systems have revolutionized the field by integrating all components into a single neural network, with attention-based encoder-decoder models achieving state-of-the-art performance. However, their autoregressive decoding process limits inference speed, making them unsuitable for real-time applications. In contrast, CTC-based models offer faster, non-autoregressive decoding but struggle to model linguistic dependencies effectively. Addressing this challenge, we propose a novel auxiliary loss framework called Language-Aware Intermediate Loss (LAIL) to enhance CTC-based ASR using the linguistic knowledge of large language models (LLMs). By attaching connector layers to intermediate encoder layers, LAIL maps outputs to the embedding space of an LLM and computes a causal language modeling loss during training. This approach enhances linguistic modeling while preserving the computational efficiency of CTC decoding. Using the Conformer architecture and various LLaMA models, we demonstrate significant improvements in Word Error Rate (WER) on the LibriSpeech, TEDLIUM2, and WSJ corpora, achieving state-of-the-art performance for CTC-based ASR with minimal computational overhead.