🤖 AI Summary
Stuttered speech—characterized by disfluencies such as blocks, prolongations, and repetitions—severely degrades automatic speech recognition (ASR) performance, leading to substantially higher word error rates (WER); further challenges include high inter- and intra-speaker variability and scarcity of annotated stuttering data, limiting model robustness and inclusivity. To address this, we conduct the first systematic comparison of personalized (speaker-specific fine-tuning) versus generalized (multi-speaker joint fine-tuning) paradigms on Whisper and Wav2Vec 2.0, incorporating multi-scenario stuttering speech augmentation and speaker-adaptive feature normalization. Results show that personalized models reduce WER by 32.7% in spontaneous speech contexts—significantly outperforming generalized models—and deliver more accurate, reliable real-time transcription in practical applications such as virtual assistants and video interviews. This work advances fairness and practical utility of ASR for people who stutter.
📝 Abstract
Stuttering -- characterized by involuntary disfluencies such as blocks, prolongations, and repetitions -- is often misinterpreted by automatic speech recognition (ASR) systems, resulting in elevated word error rates and making voice-driven technologies inaccessible to people who stutter. The variability of disfluencies across speakers and contexts further complicates ASR training, compounded by limited annotated stuttered speech data. In this paper, we investigate fine-tuning ASRs for stuttered speech, comparing generalized models (trained across multiple speakers) to personalized models tailored to individual speech characteristics. Using a diverse range of voice-AI scenarios, including virtual assistants and video interviews, we evaluate how personalization affects transcription accuracy. Our findings show that personalized ASRs significantly reduce word error rates, especially in spontaneous speech, highlighting the potential of tailored models for more inclusive voice technologies.