Seeing Speech: Learning Visible Articulatory Dynamics for Speech-Driven 3D Facial Animation

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of maintaining consistent visible articulatory motions in speech-driven 3D facial animation. To this end, it proposes an articulation-aware framework that decouples facial movements into three directional components: spreading, opening, and protrusion. A Speech-Articulation Memory (SAM) module is introduced to capture the mapping between acoustic signals and directional motions via a key-value architecture. Furthermore, a Topology-Aware Articulation Composition (TAC) algorithm is designed to facilitate feature fusion, thereby generating surface-consistent 3D animations. Experimental evaluations on benchmarks such as VOCASET demonstrate state-of-the-art performance with substantially reduced lip motion errors. User studies further confirm that the proposed approach yields superior lip-sync accuracy and perceptual realism compared to existing methods.
📝 Abstract
Recent progress in speech-driven 3D facial animation has improved vertex-level reconstruction quality, but speech-consistent visible articulation remains difficult. This is because speech production follows structured and constrained articulators' coordination and the mapping from acoustics to motion is inherently one-to-many. Motivated by the structured patterns of visible articulation, we propose a novel articulation-aware framework that models visible speech through directional articulatory motions and composes them into surface-consistent 3D facial motion. To represent visible articulation with three directional articulatory motions, spreading, opening, and protrusion, we propose a Speech--Articulatory Memory (SAM) that captures the correspondence between speech and these motions under phonetic context through retrieval and decoding based on a key-value memory structure. Then, a Topology-aware Articulatory Composition (TAC) integrates the predicted directional articulatory motions under mesh topology to produce surface-consistent 3D facial motion. Experiments on VOCASET and TFHP show that our method achieves state-of-the-art performance on standard reconstruction metrics and improves visible articulatory distance and velocity errors for lip articulation, while a user study confirms clear preference in lip sync and realism.
Problem

Research questions and friction points this paper is trying to address.

Speech-driven 3D facial animation
Visible articulatory dynamics
Acoustic-to-motion mapping
Lip sync
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speech-driven 3D facial animation
Articulatory dynamics
Key-value memory network
Topology-aware composition
Visible articulation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hyung Kyu Kim
Chung-Ang University, Seoul, South Korea
B
Byungchan Hwang
Chung-Ang University, Seoul, South Korea
Hak Gu Kim
Hak Gu Kim
Assistant Professor of GSAIM, Chung-Ang University
Machine Learning3D/AR/VR/XRRobustness of AIExplainable AI