Asymmetric Classifier-Free Guidance for Target-Speaker ASR

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the degradation of target speaker recognition performance caused by background noise and overlapping speech. Building upon the Whisper model, this work proposes an asymmetric classifier-free guidance (CFG) method. The core innovation lies in designing an asymmetric CFG decoding architecture and introducing a lightweight encoder prediction network to dynamically adjust speaker-conditioned weights, thereby achieving utterance-level adaptive guidance to optimize the decoding process. Experimental results demonstrate that the proposed method yields a 21.8% relative reduction in word error rate, significantly outperforming existing baselines. These findings indicate that the approach effectively enhances the robustness of automatic speech recognition in complex acoustic scenarios.
📝 Abstract
Target-speaker automatic speech recognition (TS-ASR) must identify and transcribe a desired speaker under varying overlap and noise conditions. These changes alter the acoustic evidence for the target speaker in the speech mixture, motivating inference-time calibration of speaker conditioning. We introduce asymmetric classifier-free guidance (CFG) for TS-ASR using Whisper: the speaker-conditioned branch predicts the target transcript, while the speaker-unconditioned branch predicts serialized multi-speaker transcripts. CFG adjusts the contribution of speaker conditioning during decoding through a single guidance scale. We select a global guidance scale on target-domain development data and train a lightweight encoder-based predictor to adjust it for each utterance, keeping the recognition model fixed. Under domain shifts, our full system achieves relative word error rate (WER) reductions of up to 21.8% over the condition-only baseline, and 5.6% over standard conditional decoding of the same CFG-trained model. Oracle analysis shows that substantially larger WER reductions are possible through utterance-level scale selection and identifies how beneficial adjustments vary with domain shifts.
Problem

Research questions and friction points this paper is trying to address.

Target-Speaker ASR
Classifier-Free Guidance
Domain Shift
Speech Recognition
Speaker Conditioning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Asymmetric Classifier-Free Guidance
Target-Speaker ASR
Whisper
Inference-Time Calibration
Utterance-Level Scale Prediction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yiwen Guan
Worcester Polytechnic Institute, Massachusetts, USA
Jacob Whitehill
Jacob Whitehill
Worcester Polytechnic Institute
Artificial Intelligence