Deep Learning Techniques for Phoneme Recognition in Italian Children's Speech

πŸ“… 2026-09-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the lack of efficient child speech-to-International Phonetic Alphabet (IPA) transcription tools for speech-language pathologists by proposing Broca, a system built upon the Conformer architecture. Employing a transfer learning strategy, the system is pretrained on adult speech and subsequently fine-tuned with fewer than three hours of pediatric data, enabling high-precision phoneme recognition in low-resource scenarios. Experimental results demonstrate that the model achieves a state-of-the-art weighted phoneme error rate of 13.36% on Italian, while exhibiting strong robustness to variations in pitch, accent, and speech disorders. This work validates the feasibility of data-efficient paradigms for clinical diagnostic assistance.
πŸ“ Abstract
Speech therapists often face difficulties diagnosing impairments due to the lack of efficient tools for transcribing speech into the International Phonetic Alphabet (IPA). This work addresses this challenge with Broca, a Conformer-based deep learning system pretrained on 8 days of adult speech and fine-tuned on a 165-minute dataset of Italian child speech collected through a range of standardized diagnostic tests for children aged 3.5-6.5. Broca was optimized to handle phonetic variability in children's speech, including tone, accent, and speech errors, and achieved a state-of-the-art weighted Phoneme Error Rate of 13.36% on Italian speech. Remarkably, this performance was obtained using less than three hours of child-specific data, underscoring the model's efficiency and robustness in low-resource clinical settings. This work demonstrates that accurate, vocabulary-independent speech-to-IPA transcription can be achieved with minimal data, paving the way for more accessible, data-efficient tools to support speech assessment and diagnosis.
Problem

Research questions and friction points this paper is trying to address.

Phoneme Recognition
Speech-to-IPA Transcription
Children's Speech
Speech Therapy
Low-resource Clinical Settings
Innovation

Methods, ideas, or system contributions that make the work stand out.

Phoneme Recognition
Conformer
Low-resource Learning
Speech-to-IPA Transcription
Children's Speech
πŸ”Ž Similar Papers
πŸ’Ό Related Jobs
No related jobs found.
N
Nicola Barbaro
University of Turin, Computer Science Department, Via Pessinetto 12, 10149 Turin, Italy
Cristina Gena
Cristina Gena
Associate professor of Computer Science, UniversitΓ  di Torino
HCIHuman Robot InteractionUser modelingHuman-centered AI
F
Francesco Petriglia
Fondazione Paideia Ente Filantropico, Via Moncalvo 1, 10131 Turin, Italy
A
Andrea Meirone
Fondazione Paideia Ente Filantropico, Via Moncalvo 1, 10131 Turin, Italy
A
Alessandro Mazzei
University of Turin, Computer Science Department, Via Pessinetto 12, 10149 Turin, Italy
A
Arianna Viotti
Fondazione Paideia Ente Filantropico, Via Moncalvo 1, 10131 Turin, Italy