FMReward: Aligning and Evaluating Audio-Driven 3D Facial Animation with Human Preferences

📅 2026-08-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the misalignment between training metrics and human perception in audio-driven 3D facial animation. We introduce FMPair, the first human preference dataset for this domain, alongside FMReward, a perception-aligned reward model, and FMFL, a direct feedback fine-tuning algorithm. By leveraging pairwise preference annotation and reward modeling, our approach enables perception-oriented optimization of diffusion models. Experimental results demonstrate that FMReward significantly outperforms traditional objective metrics in predicting human preferences. Furthermore, FMFL effectively enhances the naturalness and interactive immersion of generated animations. Collectively, this work bridges the critical gap in subjective evaluation alignment research for audio-driven 3D facial animation, establishing a new paradigm for optimizing generative models based on human perceptual feedback rather than conventional quantitative measures.
📝 Abstract
Audio-driven 3D facial animation is essential for advancing immersion and interactivity in virtual experiences. Although recent advances have shown promising capabilities, the training and evaluation of existing methods typically rely on ground-truth-based errors, which fall short of aligning with human preferences. To address this, we present a comprehensive framework that learns an automatic perceptual model from human preference data and leverages it to improve and evaluate the perceptual quality of audio-driven 3D facial animation. To begin with, we construct FMPair (Facial Motion Pairwise preference), the first human preference dataset for audio-driven 3D facial animation, which is built through a systematic annotation pipeline and comprises 65,574 annotated 3D facial motion pairs from 8,834 distinct in-the-wild audio clips. Based on the pairwise comparison dataset, we propose a Facial Motion Reward model, termed FMReward, which takes audio and 3D facial motion as inputs and predicts a perceptual quality score aligned with human preferences. Building upon FMReward, we further introduce Facial Motion reward Feedback Learning (FMFL), a direct fine-tuning algorithm that leverages a pretrained reward model to optimize diffusion-based audio-driven 3D facial animation models for better alignment with human preferences. Extensive experiments demonstrate the superiority of FMReward over other metrics in aligning with human preferences and the effectiveness of FMFL in improving the perceptual quality of audio-driven 3D facial animation.
Problem

Research questions and friction points this paper is trying to address.

Audio-driven 3D facial animation
Human preference alignment
Perceptual quality evaluation
Ground-truth metrics limitation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Human Preference Alignment
Facial Motion Reward Model
Reward Feedback Learning
Audio-Driven 3D Facial Animation
Perceptual Quality Evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Sijing Wu
Institute of Image Communication and Network Engineering, Shanghai Jiao Tong University, Shanghai, China
Y
Yunhao Li
Institute of Image Communication and Network Engineering, Shanghai Jiao Tong University, Shanghai, China
Z
Zhilin Gao
Institute of Image Communication and Network Engineering, Shanghai Jiao Tong University, Shanghai, China
Huiyu Duan
Huiyu Duan
Shanghai Jiao Tong University
Multimedia Signal Processing
Yucheng Zhu
Yucheng Zhu
Shanghai Jiaotong University
Multimedia Signal Processing
Guangtao Zhai
Guangtao Zhai
Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI EvaluationDisplays
Patrick Le Callet
Patrick Le Callet
Prof. Universite de Nantes, LS2N, Polytech Nantes - Institut Universitaire de France (IUF)
cognitive computing for MMQoEhuman perception and applications in ICT