Finding My Voice: Generative Reconstruction of Disordered Speech for Automated Clinical Evaluation

📅 2025-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Addressing the challenge of simultaneously preserving speaker identity and correcting articulation errors in automatic clinical assessment of childhood speech sound disorders (SSD), this paper proposes ChiReSSD—the first disentangled speech reconstruction framework tailored for pathological child speech. ChiReSSD explicitly decouples stylistic features (e.g., pitch, prosody) from phonemic content, integrating style-controllable text-to-speech (TTS) reconstruction, automatic phoneme recognition, and prediction of the Percent Consonants Correct (PCC) metric to jointly optimize speech repair and clinical assessability. Evaluated on the STAR dataset, ChiReSSD achieves significant improvements in word-level accuracy and speaker consistency. Its automated assessments correlate with expert ratings at ρ = 0.63, demonstrating strong cross-population generalizability. This work establishes a novel paradigm for disentangled, clinically grounded speech rehabilitation in pediatric SSD.

Technology Category

Natural Language Processing: SpeechCognitive Modeling & Cognitive Systems: Computational CreativityMachine Learning: Large Multimodal Models (LMMs)

Application Category

Search and Retrieval-Augmented AI: Assisted, interactive, and conversational searchUser Modeling, Personalization and Recommendation: User modeling and simulation for interactive and conversational systemsGraph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphs
📝 Abstract
We present ChiReSSD, a speech reconstruction framework that preserves children speaker's identity while suppressing mispronunciations. Unlike prior approaches trained on healthy adult speech, ChiReSSD adapts to the voices of children with speech sound disorders (SSD), with particular emphasis on pitch and prosody. We evaluate our method on the STAR dataset and report substantial improvements in lexical accuracy and speaker identity preservation. Furthermore, we automatically predict the phonetic content in the original and reconstructed pairs, where the proportion of corrected consonants is comparable to the percentage of correct consonants (PCC), a clinical speech assessment metric. Our experiments show Pearson correlation of 0.63 between automatic and human expert annotations, highlighting the potential to reduce the manual transcription burden. In addition, experiments on the TORGO dataset demonstrate effective generalization for reconstructing adult dysarthric speech. Our results indicate that disentangled, style-based TTS reconstruction can provide identity-preserving speech across diverse clinical populations.
Problem

Research questions and friction points this paper is trying to address.

Reconstructing disordered speech while preserving children's speaker identity
Suppressing mispronunciations in children with speech sound disorders
Automating clinical speech assessment to reduce manual transcription burden
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative reconstruction framework preserving speaker identity
Adapts to children's disordered speech emphasizing pitch
Disentangled style-based TTS for clinical populations
🔎 Similar Papers
No similar papers found.
Karen Rosero
Karen Rosero
Carnegie Mellon University
AI in HealthcareMultimodal Speech Processing
E
Eunjung Yeo
Language Technologies Institute, Carnegie Mellon University, PA, USA; Department of Computer Science, University of Texas at Austin, TX, USA
D
David R. Mortensen
Language Technologies Institute, Carnegie Mellon University, PA, USA
C
Cortney Van't Slot
Department of Plastic Surgery, University of Texas Southwestern Medical Center, TX, USA; Analytical Imaging and Modeling Center, Children’s Health, TX, USA
R
R. Hallac
Department of Plastic Surgery, University of Texas Southwestern Medical Center, TX, USA; Analytical Imaging and Modeling Center, Children’s Health, TX, USA
C
Carlos Busso
Language Technologies Institute, Carnegie Mellon University, PA, USA