MegaAvatar: Controllable Talking Avatar Generation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited global controllability over full-body pose and head motion in existing digital human generation methods. Building upon the Wan2.2-TI2V-5B model, this work introduces a pioneering latent injection mechanism based on dense SMPL-X mesh frames. By integrating a lightweight 3D convolutional encoder with an audio-visual cross-attention module, the proposed approach achieves 3D-guided full-body motion control and multimodal fine-grained expression driving, while supporting audio-driven inference without reference skeletons. The method enables high-quality talking-head video generation characterized by precise lip synchronization, robust identity consistency, and flexible spatiotemporal resolution. Consequently, it significantly enhances the global controllability of digital humans.
📝 Abstract
This report presents \textbf{MegaAvatar}, a controllable talking avatar generation framework built on top of the Wan2.2-TI2V-5B model. Compared with previous talking-avatar methods that mainly rely on audio or reference-image conditioning, we introduce additional SMPL-X-derived 3D guidance, enabling global control over body pose and head motion. Specifically, we render the driving SMPL-X sequence into dense mesh frames and encode them with a lightweight 3D convolutional encoder, whose outputs are injected into the latent tokens to provide overall motion control. Furthermore, we extend Wan2.2-TI2V-5B with additional audio and face cross-attention modules to enable fine-grained expression control and preserve the input identity, respectively. In addition, we implement an audio-to-SMPL-X model to predict an SMPL-X sequence conditioned on the reference image and input audio, allowing MegaAvatar to support audio-driven inference without user-provided SMPL-X frames. Experiments show that MegaAvatar achieves high-quality talking avatar generation with controllable body and head motion, speech-synchronized facial expressions, and consistent identity preservation. MegaAvatar also supports inference with flexible resolutions and video lengths. Codes, dataset, models will be avaliable in https://github.com/Jeoyal/MegaAvatar
Problem

Research questions and friction points this paper is trying to address.

Talking Avatar Generation
Controllable Motion
Identity Preservation
Audio-driven Synthesis
SMPL-X Guidance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Talking Avatar Generation
SMPL-X 3D Guidance
Cross-Attention Modules
Audio-driven Inference
Controllable Motion
🔎 Similar Papers
No similar papers found.