🤖 AI Summary
This study addresses the degradation of target speaker recognition performance caused by background noise and overlapping speech. Building upon the Whisper model, this work proposes an asymmetric classifier-free guidance (CFG) method. The core innovation lies in designing an asymmetric CFG decoding architecture and introducing a lightweight encoder prediction network to dynamically adjust speaker-conditioned weights, thereby achieving utterance-level adaptive guidance to optimize the decoding process. Experimental results demonstrate that the proposed method yields a 21.8% relative reduction in word error rate, significantly outperforming existing baselines. These findings indicate that the approach effectively enhances the robustness of automatic speech recognition in complex acoustic scenarios.
📝 Abstract
Target-speaker automatic speech recognition (TS-ASR) must identify and transcribe a desired speaker under varying overlap and noise conditions. These changes alter the acoustic evidence for the target speaker in the speech mixture, motivating inference-time calibration of speaker conditioning. We introduce asymmetric classifier-free guidance (CFG) for TS-ASR using Whisper: the speaker-conditioned branch predicts the target transcript, while the speaker-unconditioned branch predicts serialized multi-speaker transcripts. CFG adjusts the contribution of speaker conditioning during decoding through a single guidance scale. We select a global guidance scale on target-domain development data and train a lightweight encoder-based predictor to adjust it for each utterance, keeping the recognition model fixed. Under domain shifts, our full system achieves relative word error rate (WER) reductions of up to 21.8% over the condition-only baseline, and 5.6% over standard conditional decoding of the same CFG-trained model. Oracle analysis shows that substantially larger WER reductions are possible through utterance-level scale selection and identifies how beneficial adjustments vary with domain shifts.