🤖 AI Summary
This study addresses the challenge of recognizing speech from individuals with severe dysarthria and tracheostomies using general-purpose automatic speech recognition (ASR) systems by developing a personalized ASR framework. Methodologically, it introduces a "human-to-human dialogue" collection protocol to enhance interaction authenticity and employs acoustic simulation techniques to model pathological speech characteristics. Building upon the Whisper Base architecture, a three-stage fine-tuning strategy is designed, progressing from standard Czech to simulated pathological speech and finally to target user data. The primary contributions include the release of a novel dataset comprising 33 hours of annotated speech and a 50% reduction in character error rate compared to baseline models. The system demonstrates robust performance across scripted reading, question-answering, and spontaneous dialogue scenarios, receiving positive feedback from target users.
📝 Abstract
This work presents an automatic speech recognition (ASR) system personalized for a Czech speaker with a permanent tracheal stoma and severe dysarthria rendering their speech unintelligible to untrained listeners. We release a public dataset containing 33 annotated hours of the speaker's speech, collected using a novel "artificial conversation" protocol designed for high engagement and dialogue realism. We propose a multi-stage training pipeline based on Whisper Base: fine-tuning on standard Czech speech, acoustically simulated tracheostomic speech, and the speaker's data. We evaluate the system across three near real-time scenarios: scripted conversations, question answering, and spontaneous dialogue, achieving a 50\% relative reduction in Character Error Rate compared to Whisper Base baseline and surpassing the average recognition accuracy of their assistants in acoustic recognition of isolated utterances. We demonstrate that even for severely impeded speech, a helpful ASR is achievable, as evidenced by the quantitative results and the feedback from the speaker.