🤖 AI Summary
Existing methods struggle to balance controllability and visual fidelity: implicit representations often produce averaged expressions due to insufficient structural guidance, while explicit geometric approaches lose high-frequency textural details. To address this, this work proposes GemTalk, a novel framework that synergistically integrates implicit affective semantic features with explicit Blendshape geometric priors. Built upon a diffusion model, GemTalk employs a Video-Audio Emotion Perception (V-AEP) module to extract multimodal emotional cues, a Dynamic Geometry-Prior Generator (D-GPG) to produce identity-aware Blendshape coefficients, and a Geometry-guided Emotion Modulation (GEM) module that recalibrates the intensity of implicit features using geometric information for precise, continuous control over emotional expression. Experiments demonstrate that GemTalk significantly outperforms state-of-the-art methods in both expressive dynamics and photorealistic quality, achieving high-fidelity, highly controllable emotional talking-face generation.
📝 Abstract
Audio-driven emotional talking face generation aims to synthesize realistic videos with expressive facial dynamics. However, existing methods struggle to balance controllability and visual fidelity. Although implicit representations capture rich semantics, they lack structural guidance, often resulting in averaged emotional expressions. In contrast, explicit geometric methods offer better control over facial expressions but tend to sacrifice high-frequency texture details. To address it, we propose GemTalk, a diffusion-based framework that combines the semantic richness of implicit representations with the structural precision of explicit geometric priors. We introduce a Vision-guided Audio Emotion Projection (V-AEP) module to extract implicit emotional lip and expression features. At the same time, a Diffusion-based Geometric Priors Generator (D-GPG) generates identity-aware blendshape coefficients as explicit structural priors. Crucially, our Geometry-guided Emotion Modulation (GEM) module leverages these geometric priors to recalibrate the magnitude of implicit features, enabling precise, continuous control over emotional expressions, especially emotion intensity, without sacrificing visual quality. Extensive experiments show GemTalk achieves superior performance in photo-realism, and facial emotional dynamics.