Child-Adapted Structured Phonological Representations for Interpretable Speech Sound Analysis

📅 2026-09-30
📈 Citations: 0
✹ Influential: 0
📄 PDF
đŸ€– AI Summary
This study addresses the limitation that existing structured phonological representation models, predominantly trained on adult speech, fail to accurately capture child speech characteristics. To bridge this gap, we adapt the PhonoQ-2.0 architecture to the child speech domain by leveraging CHILDES aligned data and Montreal Forced Aligner (MFA) techniques. We systematically evaluate various supervision conditions and initialization strategies, achieving the first optimization of structured phonological representations specifically for child speech. Experimental results demonstrate that mixed adult-child forced alignment significantly outperforms purely adult-supervised approaches in manner-of-articulation recognition, attaining an F1 score of 0.987 for voicing detection. Furthermore, the proposed method effectively preserves clinically relevant longitudinal phonetic variation. This work establishes a novel paradigm for computational child speech analysis with promising implications for clinical applications.
📝 Abstract
Structured phonological representations provide an interpretable alternative to generic speech embeddings, but existing models are largely trained on adult speech. We adapt PhonoQ-2.0 to child speech using CHILDES-Aligned data and compare three alignment-supervision conditions (Adult, Adult+Child, and Child-only) across two initialization strategies (Adult PhonoQ and scratch). Generalization is evaluated against manual child-speech annotations. On 1,352 consonant targets from 58 typically developing children, child-speech adaptation improves voicing recognition across all supervision conditions, from 0.922 macro-F1 for Adult PhonoQ to 0.972--0.987 after adaptation. Manner is more sensitive to alignment supervision: Adult+Child MFA reaches 0.804 and 0.796, compared to approximately 0.70 under Adult MFA supervision. Place remains comparatively strong across systems (0.871--0.902), although per-class performance varies substantially. The velar-fronting contrast is preserved across all seven model variants. Longitudinal UltraPhonix analysis further reveals speaker-specific velar and post-alveolar changes that are largely preserved across models and broadly consistent with reported clinical progress.
Problem

Research questions and friction points this paper is trying to address.

child speech
phonological representations
speech sound analysis
interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Child-adapted phonological representations
Structured speech embeddings
Alignment supervision
Interpretable speech analysis
Longitudinal phonetic tracking
🔎 Similar Papers
No similar papers found.
A
Abner Hernandez
Pattern Recognition Lab, Friedrich-Alexander-UniversitĂ€t Erlangen-NĂŒrnberg, Germany
T
TomĂĄs Arias Vergara
Pattern Recognition Lab, Friedrich-Alexander-UniversitĂ€t Erlangen-NĂŒrnberg, Germany; GITA Lab. Facultad de IngenierĂ­a. Universidad de Antioquia UdeA, MedellĂ­n, Colombia
A
Andreas Maier
Pattern Recognition Lab, Friedrich-Alexander-UniversitĂ€t Erlangen-NĂŒrnberg, Germany
Paula Andrea Pérez-Toro
Paula Andrea Pérez-Toro
Friedrich-Alexander-UniversitĂ€t Erlangen-NĂŒrnberg; Universidad de Antioquia
Machine LearningSpeech AnalysisGait AnalysisNatural Language ProcessingDeep Learning