Cross-Learning Fine-Tuning Strategy for Dysarthric Speech Recognition Via CDSD database

πŸ“… 2025-08-26
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Dysarthric speech recognition faces significant challenges, including substantial inter-speaker variability, pronounced acoustic-phonetic deviations from healthy speech, and scarcity of speaker-specific annotated dataβ€”leading to overfitting. To address these issues, this paper proposes a cross-speaker joint fine-tuning strategy based on a pre-trained automatic speech recognition (ASR) model, simultaneously fine-tuning on dysarthric speech from multiple speakers in the CDSD corpus. Unlike conventional speaker-isolated fine-tuning paradigms, our approach leverages shared representation learning across pathological speech to enhance generalization to diverse articulatory impairments, thereby substantially reducing reliance on per-speaker labeled data. Experimental results demonstrate that the proposed method achieves up to a 13.15% absolute reduction in word error rate (WER) on target speakers compared to speaker-specific fine-tuning, with marked improvements in recognition accuracy. This work provides an efficient and practical solution for low-resource dysarthric ASR.

Technology Category

Natural Language Processing: SpeechMachine Learning: Learning with ManifoldsConstraint Satisfaction and Optimization: Distributed CSP/Optimization

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchWeb Mining and Content Analysis: Large pretrained models with web dataResponsible Web: Data and user privacy-enhancing technologies for the Web
πŸ“ Abstract
Dysarthric speech recognition faces challenges from severity variations and disparities relative to normal speech. Conventional approaches individually fine-tune ASR models pre-trained on normal speech per patient to prevent feature conflicts. Counter-intuitively, experiments reveal that multi-speaker fine-tuning (simultaneously on multiple dysarthric speakers) improves recognition of individual speech patterns. This strategy enhances generalization via broader pathological feature learning, mitigates speaker-specific overfitting, reduces per-patient data dependence, and improves target-speaker accuracy - achieving up to 13.15% lower WER versus single-speaker fine-tuning.
Problem

Research questions and friction points this paper is trying to address.

Addresses dysarthric speech recognition challenges from severity variations
Improves individual speech pattern recognition via multi-speaker fine-tuning
Reduces word error rate and data dependency through cross-learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-speaker fine-tuning for dysarthric speech
Cross-learning pathological features from multiple patients
Reducing data dependence via shared feature learning
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Q
Qing Xiao
College of Information Science and Engineering, Xinjiang University, Urumqi, China
Y
Yingshan Peng
College of Information Science and Engineering, Xinjiang University, Urumqi, China
P
PeiPei Zhang
College of Information Science and Engineering, Xinjiang University, Urumqi, China