FA-LAM: Focus-Aware Large Avatar Model for One-Shot 4D Animatable Gaussian Head

📅 2026-07-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of generating high-quality, animatable Gaussian head avatars from a single image, particularly the issues of erroneous attention activation and the inherent conflict between reconstruction fidelity and animation capability—problems that are especially pronounced in fine facial regions and under large viewing angles. To resolve these limitations, the authors propose a focus-aware large-scale head avatar model that effectively decouples reconstruction and animation objectives through semantic-symmetry attention regularization and a two-stage training pipeline. Furthermore, a visibility-aware autoregressive token fusion mechanism is introduced to enable efficient multi-view and streaming 4D reconstruction. The method significantly outperforms existing approaches in both facial detail preservation and cross-view consistency, achieving high-fidelity, fully animatable Gaussian head modeling.
📝 Abstract
We propose FA-LAM, a Focus-Aware Large Avatar Model for one-shot animatable Gaussian head creation, while simultaneously enabling static 3D and dynamic 4D full-head recovery. The core of our method lies in a thorough analysis of the attention mechanisms and the entangled reconstruction and animation training pipeline adopted by prior state-of-the-art approaches. Our analysis identifies two main factors that compromise the quality of 3D full-head generation: (1) incorrect and noisy attention activations, and (2) conflicts between the tasks of reconstruction and animation. To address the first issue, we introduce a symmetric and semantic attention regularization strategy that leverages the inherent semantics and structural symmetry of human heads. To disentangle the objectives of reconstruction and animation, we develop a novel dual-phase training pipeline that separates the model's capabilities for large-view hallucination and animation into distinct modules. Moreover, we enhance our model to support multi-view and streaming 4D reconstruction in an efficient and memory-friendly manner through a core autoregressive modification with tailored visibility-aware token fusion. Collectively, these innovations enable FA-LAM to reconstruct animatable Gaussian full heads with superior quality, particularly in fine facial regions and large viewing angles.
Problem

Research questions and friction points this paper is trying to address.

animatable avatar
3D head reconstruction
attention mechanism
one-shot learning
4D facial animation
Innovation

Methods, ideas, or system contributions that make the work stand out.

attention regularization
dual-phase training
animatable Gaussian head
4D reconstruction
visibility-aware token fusion
🔎 Similar Papers
No similar papers found.