PD-GS: Phoneme-Driven 3DGS for Audio-Driven Talking Heads

πŸ“… 2026-08-05
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the β€œleaky mouth” artifacts commonly observed in existing audio-driven 3D talking head methods, which often arise from excessive smoothing or violations of articulatory constraints such as lip closure during speech. To mitigate these issues, the authors propose a phoneme-driven Gaussian Splatting framework (PD-GS) that leverages automatic speech recognition and forced alignment to obtain temporally aligned phoneme sequences. A novel Language Fusion Module (LFM) is introduced to effectively integrate continuous audio features with discrete phoneme embeddings, thereby enhancing control over critical articulatory moments. Evaluated on the HDTF dataset, the method achieves state-of-the-art lip geometry accuracy with a Lip Movement Distance (LMD) of 2.66, significantly reducing closure violations in complex phoneme sequences and producing lip motions that exhibit greater linguistic fidelity and naturalness.
πŸ“ Abstract
3D Gaussian Splatting (3DGS) enables fast, photorealistic talking-head rendering, yet accurate lip articulation remains elusive: mouth motion is often over-smoothed and may violate hard articulatory constraints such as bilabial closures, producing the notorious ``leaky mouth'' artifact. A key difficulty is that brief, discrete articulatory events are inferred from a continuous acoustic embedding under a regression objective, which biases predictions toward averaged mouth configurations. While modern self-supervised speech encoders provide rich prosodic and phonetic cues, they do not provide an explicit, frame-aligned linguistic target that reliably disambiguates closure-level events. We propose \textbf{Phoneme-Driven Gaussian Splatting (PD-GS)}, which augments a 3DGS talker with time-aligned phoneme tokens obtained from an automatic ASR and forced-alignment pipeline. Our core component, the \textbf{Linguistic Fusion Module (LFM)}, adaptively fuses continuous audio context with discrete phoneme embeddings through a learned gate, allowing the model to preserve smooth audio-driven dynamics while strengthening phoneme guidance on articulation-critical segments. PD-GS is trained purely from monocular video using image reconstruction and lip landmark supervision. On HDTF, PD-GS achieves the best lip geometry among the compared baselines (LMD 2.66) and qualitatively reduces closure violations in challenging phoneme sequences, yielding more linguistically faithful neural avatars.
Problem

Research questions and friction points this paper is trying to address.

audio-driven talking heads
lip articulation
phoneme alignment
3D Gaussian Splatting
articulatory constraints
Innovation

Methods, ideas, or system contributions that make the work stand out.

Phoneme-Driven
3D Gaussian Splatting
Linguistic Fusion Module
Audio-Driven Talking Heads
Forced Alignment
πŸ”Ž Similar Papers
2024-03-19IEEE Workshop/Winter Conference on Applications of Computer VisionCitations: 4