🤖 AI Summary
This study addresses the challenge of early screening for sibilant substitution disorders in Polish-speaking children under conditions of limited clinical resources by proposing a lightweight, home-deployable conservative screening protocol. The approach integrates a wav2vec2 CTC-based phoneme recognizer with forced alignment to detect mispronunciations and employs template-driven natural language generation combined with rule-based tagging to deliver interpretable feedback to caregivers. Evaluated on 559 previously unseen child utterances, the system achieves 88.7% sequence-level matching accuracy; for the target substitution errors, it attains 72.9% precision and 61.4% recall (F1 = 0.67), with a low false alarm rate of 2.7%. A clinician-in-the-loop validation mechanism is incorporated to ensure safety margins in diagnostic decision-making.
📝 Abstract
Early identification of speech sound errors in children is often limited by access to specialists, motivating lightweight screening tools that can operate outside the clinic. We present a screening pipeline for Polish-speaking children focused on sibilant substitutions, coupling a wav2vec2-based CTC token recognizer with alignment-based error typing and a template-grounded caregiver assistant for screening, not diagnosis. On a held-out test set of 10 unseen children comprising 559 utterances, the recognizer achieves 88.7 percent exact sequence match. As a conservative screening proxy, we flag a mismatch when the system emits substitution-evidence bracketed tokens at the target segment, yielding 72.9 percent precision, 61.4 percent recall, F1 = 0.67, and a 2.7 percent false-alarm rate on target-correct items. We describe the assistant's safety boundaries and outline a clinician-in-the-loop validation plan for future deployment.