Boosting CTC-Based ASR Using LLM-Based Intermediate Loss Regularization

📅 2025-06-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
CTC-based ASR models offer efficient non-autoregressive decoding but suffer from weak language modeling capability, limiting recognition accuracy. To address this, we propose an LLM-driven intermediate loss regularization framework: leveraging the LLaMA embedding space as semantic supervision, we map intermediate representations from a Conformer encoder into the LLM’s semantic space via a learnable projection layer and jointly optimize with a causal language modeling loss. Crucially, this approach enhances speech features with linguistic knowledge without modifying the CTC decoding pipeline. Experiments on LibriSpeech, TEDLIUM2, and WSJ demonstrate substantial WER reductions, achieving state-of-the-art performance among CTC-based methods with negligible computational overhead. Our key contribution is the first use of an LLM’s semantic space as an intermediate supervision signal, enabling lightweight, efficient, and end-to-end language-aware regularization.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsPlanning, Routing, and Scheduling: Planning with Language Models

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
End-to-end (E2E) automatic speech recognition (ASR) systems have revolutionized the field by integrating all components into a single neural network, with attention-based encoder-decoder models achieving state-of-the-art performance. However, their autoregressive decoding process limits inference speed, making them unsuitable for real-time applications. In contrast, CTC-based models offer faster, non-autoregressive decoding but struggle to model linguistic dependencies effectively. Addressing this challenge, we propose a novel auxiliary loss framework called Language-Aware Intermediate Loss (LAIL) to enhance CTC-based ASR using the linguistic knowledge of large language models (LLMs). By attaching connector layers to intermediate encoder layers, LAIL maps outputs to the embedding space of an LLM and computes a causal language modeling loss during training. This approach enhances linguistic modeling while preserving the computational efficiency of CTC decoding. Using the Conformer architecture and various LLaMA models, we demonstrate significant improvements in Word Error Rate (WER) on the LibriSpeech, TEDLIUM2, and WSJ corpora, achieving state-of-the-art performance for CTC-based ASR with minimal computational overhead.
Problem

Research questions and friction points this paper is trying to address.

Improving linguistic dependencies in CTC-based ASR models
Enhancing ASR performance without sacrificing decoding speed
Integrating LLM knowledge to boost CTC model accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-based intermediate loss regularization
Connector layers for embedding space mapping
Causal language modeling loss integration
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Independent Researcher
D
Duygu Altinok
Independent Researcher, Germany