🤖 AI Summary
This study addresses the challenge of maintaining consistent visible articulatory motions in speech-driven 3D facial animation. To this end, it proposes an articulation-aware framework that decouples facial movements into three directional components: spreading, opening, and protrusion. A Speech-Articulation Memory (SAM) module is introduced to capture the mapping between acoustic signals and directional motions via a key-value architecture. Furthermore, a Topology-Aware Articulation Composition (TAC) algorithm is designed to facilitate feature fusion, thereby generating surface-consistent 3D animations. Experimental evaluations on benchmarks such as VOCASET demonstrate state-of-the-art performance with substantially reduced lip motion errors. User studies further confirm that the proposed approach yields superior lip-sync accuracy and perceptual realism compared to existing methods.
📝 Abstract
Recent progress in speech-driven 3D facial animation has improved vertex-level reconstruction quality, but speech-consistent visible articulation remains difficult. This is because speech production follows structured and constrained articulators' coordination and the mapping from acoustics to motion is inherently one-to-many. Motivated by the structured patterns of visible articulation, we propose a novel articulation-aware framework that models visible speech through directional articulatory motions and composes them into surface-consistent 3D facial motion. To represent visible articulation with three directional articulatory motions, spreading, opening, and protrusion, we propose a Speech--Articulatory Memory (SAM) that captures the correspondence between speech and these motions under phonetic context through retrieval and decoding based on a key-value memory structure. Then, a Topology-aware Articulatory Composition (TAC) integrates the predicted directional articulatory motions under mesh topology to produce surface-consistent 3D facial motion. Experiments on VOCASET and TFHP show that our method achieves state-of-the-art performance on standard reconstruction metrics and improves visible articulatory distance and velocity errors for lip articulation, while a user study confirms clear preference in lip sync and realism.