Expression-Diverse References for Identity-Preserving Video Generation

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the identity distortion and evaluation bias in single-reference identity-preserving video generation caused by expression variations. We propose a hierarchical evaluation framework based on expression intensity alongside an automated data curation pipeline. Specifically, our method leverages pretrained facial reenactment models to synthesize missing expression samples, thereby constructing an expression-diverse reference set. The selection process is further optimized using a Stand-In architecture, substantially enhancing identity consistency under extreme expressions. Experimental results demonstrate that, across both real and synthetic reference conditions, the proposed approach achieves superior identity similarity compared to baseline methods within all expression intensity intervals, with particularly pronounced improvements for extreme expressions. This work provides an effective paradigm for robust identity-preserving video generation.
📝 Abstract
Identity-preserving video generation aims to maintain a subject's identity while synthesizing realistic videos. Yet a single reference portrait captures the subject's appearance under only one facial configuration. As expressions change, facial appearance can vary in highly identity-specific ways, leaving the subject's appearance under unseen expressions underdetermined by the reference alone. This expression-dependent variation also complicates evaluation: similarity to a neutral reference may decrease under strong expressions even for real images of the same person. We investigate this limitation from both generation and evaluation perspectives. First, we quantify how face-recognition similarity varies with expression intensity using controlled photographs and MEAD videos. We then construct a compact yet expressive reference gallery that captures diverse expression-dependent facial configurations. Matching against this gallery provides a more robust measure of identity similarity under expressive motion. To further expose performance degradation with expression intensity, we report identity similarity separately for mild, intense, and extreme expressions. For generation, we extend Stand-In to condition on our expression-diverse reference sets and develop a data-curation pipeline that extracts consistent yet diverse face crops from training videos. In practical settings where only a single portrait is available, we construct the reference set by synthesizing additional expressions with a pretrained facial reenactment model. On our controlled benchmark, both real and synthesized reference sets outperform the evaluated baselines in identity similarity across all three expression-intensity regimes, with the largest improvements for extreme expressions.
Problem

Research questions and friction points this paper is trying to address.

Identity-preserving video generation
Expression diversity
Identity similarity evaluation
Facial expression variation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Identity-Preserving Video Generation
Expression-Diverse References
Facial Reenactment
Data-Curation Pipeline
Expression-Dependent Evaluation
T
Tianwen Fu
Institute for Creative Technologies, University of Southern California
Wenbin Teng
Wenbin Teng
University of Southern California
Computer VisionGenerative Model3D reconstruction
Gonglin Chen
Gonglin Chen
University of Southern California
Computer VisionMachine Learning
J
Junyi Ouyang
Institute for Creative Technologies, University of Southern California
H
Haolin Xiong
Institute for Creative Technologies, University of Southern California
Yajie Zhao
Yajie Zhao
Computer Scientist at University of Southern California
Virtual HumanNeural RenderAR/VR