🤖 AI Summary
Addressing the challenge of simultaneously preserving speaker identity and correcting articulation errors in automatic clinical assessment of childhood speech sound disorders (SSD), this paper proposes ChiReSSD—the first disentangled speech reconstruction framework tailored for pathological child speech. ChiReSSD explicitly decouples stylistic features (e.g., pitch, prosody) from phonemic content, integrating style-controllable text-to-speech (TTS) reconstruction, automatic phoneme recognition, and prediction of the Percent Consonants Correct (PCC) metric to jointly optimize speech repair and clinical assessability. Evaluated on the STAR dataset, ChiReSSD achieves significant improvements in word-level accuracy and speaker consistency. Its automated assessments correlate with expert ratings at ρ = 0.63, demonstrating strong cross-population generalizability. This work establishes a novel paradigm for disentangled, clinically grounded speech rehabilitation in pediatric SSD.
📝 Abstract
We present ChiReSSD, a speech reconstruction framework that preserves children speaker's identity while suppressing mispronunciations. Unlike prior approaches trained on healthy adult speech, ChiReSSD adapts to the voices of children with speech sound disorders (SSD), with particular emphasis on pitch and prosody. We evaluate our method on the STAR dataset and report substantial improvements in lexical accuracy and speaker identity preservation. Furthermore, we automatically predict the phonetic content in the original and reconstructed pairs, where the proportion of corrected consonants is comparable to the percentage of correct consonants (PCC), a clinical speech assessment metric. Our experiments show Pearson correlation of 0.63 between automatic and human expert annotations, highlighting the potential to reduce the manual transcription burden. In addition, experiments on the TORGO dataset demonstrate effective generalization for reconstructing adult dysarthric speech. Our results indicate that disentangled, style-based TTS reconstruction can provide identity-preserving speech across diverse clinical populations.