Re-Sonance: A Dysarthric Asynchronous Real-Time Speech Conversion System Based on a Three-Stage Cascaded ASR-LLM-TTS Architecture

📅 2026-07-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge faced by individuals with dysarthria in professional settings, where existing augmentative and alternative communication systems often suffer from high latency and unnatural speech output, hindering real-time interaction. To overcome these limitations, the authors propose a real-time speech conversion system based on a cascaded architecture integrating automatic speech recognition (ASR), large language models (LLMs), and text-to-speech (TTS). This approach uniquely incorporates an LLM into the dysarthric speech conversion pipeline, leveraging Whisper for ASR, Qwen for semantic enhancement, and CosyVoice for high-quality speech synthesis. The system achieves low latency while significantly improving intelligibility, naturalness, and semantic coherence of the generated speech. Experimental results on a Chinese dysarthric speech dataset demonstrate superior performance in both subjective and objective evaluations, offering an effective and practical communication solution for individuals with mild to moderate dysarthria.
📝 Abstract
Individuals with dysarthria face significant challenges in professional speaking scenarios such as conferences, presentations, and meetings, where real-time communication is crucial. While existing Augmentative and Alternative Communication (AAC) systems provide basic support, they often fail to meet the demands of professional speaking environments due to high latency and unnatural speech patterns. This paper presents Re-Sonance, a novel LLM-enhanced speech-driven AAC system designed for real-time professional speaking scenarios. By integrating Whisper ASR, Qwen LLM, and CosyVoice TTS, Re-Sonance achieves improved speech intelligibility and naturalness while maintaining real-time performance. Both subjective and objective evaluations using a Mandarin dysarthric speech dataset demonstrate that our speech reconstruction approach significantly improved intelligibility while preserving semantic coherence for speakers with mild to moderate dysarthria. Although performance remains limited for severe dysarthria cases, our findings validate the potential of LLM-based methods for enhancing speech-driven AAC systems, paving the way for more effective and accessible communication technologies.
Problem

Research questions and friction points this paper is trying to address.

dysarthria
real-time speech conversion
Augmentative and Alternative Communication (AAC)
speech intelligibility
professional speaking scenarios
Innovation

Methods, ideas, or system contributions that make the work stand out.

dysarthric speech conversion
real-time AAC
ASR-LLM-TTS cascade
speech intelligibility
large language model