Customizing Speech Recognition Model with Large Language Model Feedback

📅 2025-06-05
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
ASR systems exhibit significant performance degradation in recognizing rare named entities and adapting to out-of-domain speech. To address this, we propose an unsupervised end-to-end domain adaptation method that integrates large language models (LLMs) as context-aware reward models within a reinforcement learning framework—specifically, Proximal Policy Optimization (PPO)—to generate label-free feedback signals for ASR model fine-tuning. Unlike conventional approaches relying on manual annotations or pseudo-labeling, our method leverages LLMs to dynamically assess the semantic plausibility of ASR hypotheses, thereby guiding policy optimization. On named entity recognition benchmarks, our approach reduces word error rate by 21% compared to standard self-training baselines, markedly improving cross-domain robustness. The core contribution lies in the native integration of LLMs into the ASR RL pipeline, enabling fully unsupervised, context-sensitive, and semantics-driven domain adaptation without any human supervision or external labeling.

Technology Category

Machine Learning: Transfer, Domain Adaptation, Multi-Task LearningNatural Language Processing: (Large) Language ModelsSearch and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Vertical and domain-specific searchUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Automatic speech recognition (ASR) systems have achieved strong performance on general transcription tasks. However, they continue to struggle with recognizing rare named entities and adapting to domain mismatches. In contrast, large language models (LLMs), trained on massive internet-scale datasets, are often more effective across a wide range of domains. In this work, we propose a reinforcement learning based approach for unsupervised domain adaptation, leveraging unlabeled data to enhance transcription quality, particularly the named entities affected by domain mismatch, through feedback from a LLM. Given contextual information, our framework employs a LLM as the reward model to score the hypotheses from the ASR model. These scores serve as reward signals to fine-tune the ASR model via reinforcement learning. Our method achieves a 21% improvement on entity word error rate over conventional self-training methods.
Problem

Research questions and friction points this paper is trying to address.

Improving recognition of rare named entities in ASR
Adapting ASR models to domain mismatches unsupervised
Enhancing transcription quality using LLM feedback signals
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement learning fine-tunes ASR model
LLM feedback scores ASR hypotheses as reward
Unsupervised adaptation using unlabeled data context
🔎 Similar Papers
No similar papers found.
S
Shaoshi Ling
Microsoft Core AI, Redmond, WA, USA
G
Guoli Ye
Microsoft Core AI, Redmond, WA, USA