Fine-Tuning ASR for Stuttered Speech: Personalized vs. Generalized Approaches

📅 2025-06-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Stuttered speech—characterized by disfluencies such as blocks, prolongations, and repetitions—severely degrades automatic speech recognition (ASR) performance, leading to substantially higher word error rates (WER); further challenges include high inter- and intra-speaker variability and scarcity of annotated stuttering data, limiting model robustness and inclusivity. To address this, we conduct the first systematic comparison of personalized (speaker-specific fine-tuning) versus generalized (multi-speaker joint fine-tuning) paradigms on Whisper and Wav2Vec 2.0, incorporating multi-scenario stuttering speech augmentation and speaker-adaptive feature normalization. Results show that personalized models reduce WER by 32.7% in spontaneous speech contexts—significantly outperforming generalized models—and deliver more accurate, reliable real-time transcription in practical applications such as virtual assistants and video interviews. This work advances fairness and practical utility of ASR for people who stutter.

Technology Category

Natural Language Processing: SpeechMachine Learning: Large Multimodal Models (LMMs)Cognitive Modeling & Cognitive Systems: Adaptive Behavior

Application Category

Search and Retrieval-Augmented AI: Personalized, context-aware and across-device searchUser Modeling, Personalization and Recommendation: User privacy protection in personalized systemsResponsible Web: Data and user privacy-enhancing technologies for the Web
📝 Abstract
Stuttering -- characterized by involuntary disfluencies such as blocks, prolongations, and repetitions -- is often misinterpreted by automatic speech recognition (ASR) systems, resulting in elevated word error rates and making voice-driven technologies inaccessible to people who stutter. The variability of disfluencies across speakers and contexts further complicates ASR training, compounded by limited annotated stuttered speech data. In this paper, we investigate fine-tuning ASRs for stuttered speech, comparing generalized models (trained across multiple speakers) to personalized models tailored to individual speech characteristics. Using a diverse range of voice-AI scenarios, including virtual assistants and video interviews, we evaluate how personalization affects transcription accuracy. Our findings show that personalized ASRs significantly reduce word error rates, especially in spontaneous speech, highlighting the potential of tailored models for more inclusive voice technologies.
Problem

Research questions and friction points this paper is trying to address.

ASR systems misinterpret stuttered speech, increasing word errors
Limited annotated stuttered speech data complicates ASR training
Personalized vs. generalized ASR models for stuttering accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fine-tuning ASR for stuttered speech
Comparing personalized vs. generalized models
Personalized ASRs reduce word error rates
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Dena F. Mujtaba
Michigan State University, USA
N
N. Mahapatra
Michigan State University, USA