🤖 AI Summary
This study addresses the tendency of existing speech-driven facial animation models to oversimplify lip transitions, resulting in lost coarticulation trajectories and distorted articulatory movements. To tackle this issue, we introduce a novel ground-truth-free geometric metric based on lip path length and systematically evaluate four mainstream models using forced alignment and a large-scale user perception study. Our findings reveal that current models universally suffer from trajectory flattening, missing 15%–60% of rapid articulatory components, while viewers significantly prefer authentic speech trajectories. This work is the first to quantitatively expose this shared bottleneck in the field, establishing clear improvement targets and an evaluation benchmark for high-fidelity speech-driven animation.
📝 Abstract
Speech-driven 3D facial animation can reproduce recognizable mouth poses. However, it can simplify the motion between them, and that motion carries coarticulation, the way the sounds around each sound shape its articulation. We introduce a geometric measure of this trajectory shaping: lip-path length compared with the shortest route through the vowel, consonant and vowel positions of a speech segment. In contrast to the endpoint chord, this consonant-aware route accounts for obligatory transit and avoids degeneracy, while preserving invariance to uniform motion gain. The measure needs only a forced alignment, so it applies where no ground truth exists. We demonstrate it on four state-of-the-art methods, one per architectural family, real-time and offline. All four trace flatter lip trajectories than captured speech. Against frame-rate-matched ground truth, DiffPoseTalk, ARTalk and FaceFormer show clear deficits, equivalent on this measure to removing 15-60% of real speech's fast articulatory component. CodeTalker is marginal on the primary measure and clear on a companion measure. A pre-registered study with 97 viewers and 3,523 judgments underpins the measured direction: controlled damping of real motion lowers the score and is penalized, whereas exaggeration shows no detected penalty over the tested range. Viewers also prefer real speech in 73.4% of sentence comparisons and, in the aggregate, on single words. Together, the measure, its calibration and the study identify a perceptually relevant loss of trajectory shaping and a concrete target for improving synthesized articulation.