Phoneme-Level Mispronunciation Screening in Polish-Speaking Children with an Explainable Assistant

📅 2026-06-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of early screening for sibilant substitution disorders in Polish-speaking children under conditions of limited clinical resources by proposing a lightweight, home-deployable conservative screening protocol. The approach integrates a wav2vec2 CTC-based phoneme recognizer with forced alignment to detect mispronunciations and employs template-driven natural language generation combined with rule-based tagging to deliver interpretable feedback to caregivers. Evaluated on 559 previously unseen child utterances, the system achieves 88.7% sequence-level matching accuracy; for the target substitution errors, it attains 72.9% precision and 61.4% recall (F1 = 0.67), with a low false alarm rate of 2.7%. A clinician-in-the-loop validation mechanism is incorporated to ensure safety margins in diagnostic decision-making.
📝 Abstract
Early identification of speech sound errors in children is often limited by access to specialists, motivating lightweight screening tools that can operate outside the clinic. We present a screening pipeline for Polish-speaking children focused on sibilant substitutions, coupling a wav2vec2-based CTC token recognizer with alignment-based error typing and a template-grounded caregiver assistant for screening, not diagnosis. On a held-out test set of 10 unseen children comprising 559 utterances, the recognizer achieves 88.7 percent exact sequence match. As a conservative screening proxy, we flag a mismatch when the system emits substitution-evidence bracketed tokens at the target segment, yielding 72.9 percent precision, 61.4 percent recall, F1 = 0.67, and a 2.7 percent false-alarm rate on target-correct items. We describe the assistant's safety boundaries and outline a clinician-in-the-loop validation plan for future deployment.
Problem

Research questions and friction points this paper is trying to address.

mispronunciation screening
speech sound errors
sibilant substitutions
children
early identification
Innovation

Methods, ideas, or system contributions that make the work stand out.

phoneme-level mispronunciation
wav2vec2-based CTC
explainable assistant
sibilant substitution
alignment-based error typing
🔎 Similar Papers
No similar papers found.
M
Milosz Dudek
AGH University of Krakow, Cracow, Poland
Daria Hemmerling
Daria Hemmerling
AGH University of Science and Technology, Department of Measurement and Electronics // SoftServe
Biomedical EngineeringSignal ProcessingMachine Learning
K
Kamil Kwarciak
AGH University of Krakow, Cracow, Poland
M
Maciej Stroinski
SoftServe, Cracow, Poland
M
Maria Pensko
SoftServe, Cracow, Poland
M
Mateusz Kowalewski
SoftServe, Cracow, Poland
L
Leonid Pavlovskyi
SoftServe, Cracow, Poland
S
Sebastian Jurczak
SoftServe, Cracow, Poland
A
Anna-Mariia Vitkovska
SoftServe, Cracow, Poland
Z
Zuzanna Miodonska
Department of Biomedical Engineering, Silesian University of Technology, Poland
N
Natalia Mocko
Institute of Linguistics, Faculty of Humanities, University of Silesia in Katowice, Poland
M
Michal Krecichwost
Department of Biomedical Engineering, Silesian University of Technology, Poland