EvolvingAvatar: Interactive 3D Head Generation That Adapts as Conversations Unfold

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing 3D head generators with fixed parameters, which struggle to leverage conversational patterns as learning signals. To this end, we propose a causal generator framework that dynamically adapts to audiovisual contexts through test-time training, enabling interactive 3D head generation. The method introduces a binary context prediction objective to provide self-supervised signals and designs persistent fast weights alongside transient jaw adaptation mechanisms to effectively integrate short- and long-term conversational features. Furthermore, we construct InterHead-Bench, a dedicated evaluation benchmark. Experimental results demonstrate that our approach reduces expression statistical mismatch by 11.1% on out-of-distribution data, significantly enhancing the quality of conversational motion.
📝 Abstract
Interactive 3D head generation requires coordinated speaking and listening motion that responds to an evolving conversation. Existing generators use incoming observations as context but keep their parameters fixed, leaving conversational patterns unused as a learning signal. We introduce EvolvingAvatar, a causal generator that uses test-time training to adapt to user face video and dyadic audio during interaction. Its dyadic context prediction objective provides a self-supervised learning signal from audiovisual context without target motion labels at test time. Persistent fast weights accumulate these updates within each conversation to guide motion generation, while transient jaw adaptation responds to current audiovisual context. Predicted speech activity controls how persistent adaptation guides motion. We also introduce InterHead-Bench, a unified 455.95-hour benchmark built from single-view and dual-view conversation videos. Experiments show improved conversational motion statistics over strong baselines. On the hardest out-of-distribution split, generation improves as conversations unfold, reducing mismatch with recorded user-avatar expression statistics by up to 11.1% from the first interval.
Problem

Research questions and friction points this paper is trying to address.

Interactive 3D head generation
Conversational motion
Test-time adaptation
Dyadic interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Time Training
Self-Supervised Learning
Causal Generator
Persistent Fast Weights
Interactive 3D Head Generation
🔎 Similar Papers
No similar papers found.
J
Junjie Chen
Hefei University of Technology; EPIC Lab, Shanghai Jiao Tong University
F
Fei Wang
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center
K
Kun Li
United Arab Emirates University
Y
Yiqi Nie
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center; Anhui University
X
Xun Yang
University of Science and Technology of China
Yanbin Hao
Yanbin Hao
Hefei University of Technology
Video retrievalvideo action recognitionhashingVideo Hyperlinking
Linfeng Zhang
Linfeng Zhang
DP Technology; AI for Science Institute
AI for Sciencemulti-scale modelingmolecular simulationdrug/materials design
M
Meng Wang
Hefei University of Technology