MagicFace: High-Fidelity Facial Expression Editing with Action-Unit Control

📅 2025-01-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses high-fidelity, controllable single-subject facial expression editing. Methodologically, it introduces a novel action unit (AU) delta-driven conditional diffusion model. Specifically, it proposes the first diffusion framework explicitly conditioned on relative AU intensity changes; designs a self-attention-based identity encoder to preserve subject-specific appearance; incorporates explicit pose and background attribute controllers to ensure geometric and scene consistency; and embeds continuous, interpretable AU delta signals into the UNet denoising process. Experiments demonstrate that the method generates natural, identity-preserving expression animations under multi-AU combinations, significantly outperforming state-of-the-art approaches in both qualitative and quantitative evaluations. It supports arbitrary identity transfer while maintaining facial fidelity and temporal coherence. The implementation is publicly available.

Technology Category

Computer Vision: Diffusion Models for VisionHumans and AI: Human-Aware Planning and Behavior PredictionCognitive Modeling & Cognitive Systems: Affective Computing

Application Category

User Modeling, Personalization and Recommendation: Accountability, Transparency, and Ethics for personalizationEconomics, Online Markets and Human Computation: LLM based quality controls for crowd workSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
We address the problem of facial expression editing by controling the relative variation of facial action-unit (AU) from the same person. This enables us to edit this specific person's expression in a fine-grained, continuous and interpretable manner, while preserving their identity, pose, background and detailed facial attributes. Key to our model, which we dub MagicFace, is a diffusion model conditioned on AU variations and an ID encoder to preserve facial details of high consistency. Specifically, to preserve the facial details with the input identity, we leverage the power of pretrained Stable-Diffusion models and design an ID encoder to merge appearance features through self-attention. To keep background and pose consistency, we introduce an efficient Attribute Controller by explicitly informing the model of current background and pose of the target. By injecting AU variations into a denoising UNet, our model can animate arbitrary identities with various AU combinations, yielding superior results in high-fidelity expression editing compared to other facial expression editing works. Code is publicly available at https://github.com/weimengting/MagicFace.
Problem

Research questions and friction points this paper is trying to address.

Facial Expression Modification
Precise Facial Action Units Adjustment
Identity Preservation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Advanced Facial Expression Editing
UNet Model Integration
Expression Modification Precision
🔎 Similar Papers
No similar papers found.