Optimizing Multi-Stuttered Speech Classification: Leveraging Whisper's Encoder for Efficient Parameter Reduction in Automated Assessment

📅 2024-06-09
🏛️ arXiv.org
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the automatic classification of multi-repetitive stuttered speech. We introduce the first annotated dataset specifically designed for diverse stuttering disfluency types and propose a lightweight multi-label classification framework built upon the Whisper encoder. A key insight is that the final Whisper encoder layer exhibits the highest discriminative power for stutter detection; leveraging this, we devise a hierarchical freezing strategy—fine-tuning only a single layer—achieving high performance while drastically reducing model size from 20.27M to 3.29M parameters (83.7% compression). This yields significantly improved cross-dialect and cross-lingual adaptability. Evaluated on the Fluency-Bank test set, our method achieves micro-, macro-, and weighted F1-scores of 0.88, 0.85, and 0.87, respectively, demonstrating both effectiveness and strong generalization capability.

Technology Category

Natural Language Processing: SpeechMachine Learning: Large Multimodal Models (LMMs)Computer Vision: Multi-modal Vision

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchWeb Mining and Content Analysis: Large pretrained models with web dataResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
The automated classification of stuttered speech has significant implications for timely assessments providing assistance to speech language pathologists. Despite notable advancements in the field, the cases in which multiple disfluencies occur in speech require attention. We have taken a progressive approach to fill this gap by classifying multi-stuttered speech more efficiently. The problem has been addressed by firstly curating a dataset of multi-stuttered disfluencies from open source dataset SEP-28k audio clips. Secondly, employing Whisper, a state-of-the-art speech recognition model has been leveraged by using its encoder and taking the problem as multi label classification. Thirdly, using a 6 encoder layer Whisper and experimenting with various layer freezing strategies, a computationally efficient configuration of the model was identified. The proposed configuration achieved micro, macro, and weighted F1-scores of 0.88, 0.85, and 0.87, correspondingly on an external test dataset i.e. Fluency-Bank. In addition, through layer freezing strategies, we were able to achieve the aforementioned results by fine-tuning a single encoder layer, consequently, reducing the model's trainable parameters from 20.27 million to 3.29 million. This research study unveils the contribution of the last encoder layer in the identification of disfluencies in stuttered speech. Consequently, it has led to a computationally efficient approach, 83.7% less parameters to train, making the proposed approach more adaptable for various dialects and languages.
Problem

Research questions and friction points this paper is trying to address.

Classifying multi-stuttered speech efficiently
Reducing model parameters for computational efficiency
Leveraging Whisper's encoder for speech disfluencies identification
Innovation

Methods, ideas, or system contributions that make the work stand out.

Utilizes Whisper's encoder
Implements layer freezing
Reduces trainable parameters
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Huma Ameer
Huma Ameer
S
Seemab Latif
R
R. Latif