🤖 AI Summary
This study addresses the hierarchical mapping problem from abstract linguistic representations to concrete articulatory–acoustic realizations in speech production. To this end, we propose DYNARTmo—a dynamic articulatory model grounded in a neurobiologically inspired framework. DYNARTmo employs speech gestures and their score-like temporal representations (“gesture scores”) as control signals, driving a vocal-tract model via gesture coordination mechanisms and continuous trajectory generation to yield physiologically plausible articulator movements. It realizes an end-to-end, hierarchical computational simulation: language → gestures → articulation → acoustics. Key contributions include: (1) the first integration of gesture scores into dynamic articulatory modeling, explicitly encoding temporal coordination; (2) real-time, continuous articulatory synthesis for complex phonetic sequences; and (3) differentiable, biophysically constrained joint simulation of acoustic and articulatory parameters. Experiments demonstrate high-fidelity modeling of multisyllabic words and connected speech.
📝 Abstract
This paper describes the current implementation of the dynamic articulatory model DYNARTmo, which generates continuous articulator movements based on the concept of speech gestures and a corresponding gesture score. The model provides a neurobiologically inspired computational framework for simulating the hierarchical control of speech production from linguistic representation to articulatory-acoustic realization. We present the structure of the gesture inventory, the coordination of gestures in the gesture score, and their translation into continuous articulator trajectories controlling the DYNARTmo vocal tract model.