🤖 AI Summary
This study addresses the challenges faced by large speech models in efficiently utilizing large-scale biasing lists and recognizing rare words. To this end, we propose a two-stage contextual biasing framework that introduces a novel phoneme-level temporal competition mechanism. By sharing phoneme posteriors, this approach enables bias retrieval without additional forward passes. Furthermore, it integrates frame-synchronous decoding with local correction for post-decoding optimization, effectively balancing computational efficiency and recognition accuracy. Experimental results on the LibriSpeech dataset demonstrate that the proposed method significantly reduces the bias word error rate (B-WER) by over 23% while maintaining stable general word error rates. This work thus provides an effective solution to the problem of rare word recognition in large-scale biasing scenarios.
📝 Abstract
Contextual biasing improves rare-word recognition in speech large language models (SpeechLLMs), but efficiently exploiting large bias lists remains challenging. We propose PTC-Bias, a two-stage framework based on phoneme-level temporal competition. At the prefill stage, PTC Retrieval performs frame-synchronous phoneme decoding and temporal competition among candidate pronunciations, producing a compact bias-word shortlist and corresponding speech intervals. After SpeechLLM decoding, PTC Correction conducts a second local competition between the retrieved candidates and mismatched transcript spans within these intervals. Selective correction reduces near-homophone and word-segmentation errors while preserving correct transcriptions. Both stages share the same phoneme posteriors and require no additional SpeechLLM forward pass. Experiments on LibriSpeech show consistent gains across two SpeechLLMs and bias lists of up to 2000 words. With Prompt-SLAM-ASR-7B and 2000 bias words, PTC-Bias reduces B-WER by 23.4%/23.9% relative to CTC-Filter on test-clean/test-other, while keeping U-WER nearly unchanged.