🤖 AI Summary
This study addresses the limitations of existing talking head generation methods, which often overlook individual speaking habits and produce monotonous facial motions. To this end, we propose a personalized, real-time talking head generation approach that captures speaker-specific characteristics through motion space modeling and two-stage imitation learning. Furthermore, a single-step flow matching framework is introduced to enable efficient inference. We also design the PLAD metric to precisely quantify subtle differences in articulatory habits. By integrating flow matching, imitation learning, and motion projection techniques, the proposed method significantly improves the fidelity of speaking habit imitation, achieving high-quality, real-time digital human generation.
📝 Abstract
In daily life, each person exhibits unique speaking habits, leading to subtle yet consistent lip-shape variations even when pronouncing the same word. Although recent talking head generation methods have achieved impressive visual fidelity and lip synchronization, they largely overlook user-specific customization, especially the motion patterns that characterize individual speaking habits. These habits are difficult to model and capture, as their motion patterns are highly fine-grained and often similar across individuals. As a result, many approaches produce overly uniform facial motions and fail to capture diverse, person-specific articulation patterns. To address this, we propose TalkLikeYou, an efficient framework that imitates how a target person speaks in talking head generation. Our method models habit in motion-space and achieves real-time performance through Flow Matching with only one sampling step during inference. We further adopt a two-stage imitation learning strategy to capture subtle distinctions between habits, allowing users to specify a target habit through either a preset style from the dataset or a reference video. In addition, we introduce a new metric PLAD that projects mouth motions onto representative articulation axes to evaluate imitation accuracy and generation diversity. Extensive experiments demonstrate that TalkLikeYou generates high-quality talking heads in real-time and significantly improves speaking habit imitation compared with prior methods. The code is available at: https://github.com/BQ-Wang0511/TalkLikeYou