GHARP: Real-time Gaussian Head Animation from Large-scale Reconstruction Prior

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of identity-expression coupling, torso region ambiguity, and high computational overhead in real-time 3D head animation by proposing a two-stage framework comprising offline identity construction and online lightweight animation. Methodologically, it leverages a semantically structured latent space derived from pretrained reconstruction priors to predict expression residuals, thereby achieving effective decoupling. Additionally, a body-alignment network is designed to eliminate training ambiguities, while 3D Gaussian Splatting is integrated for efficient rendering. Evaluated on the Ava-256 benchmark, the proposed framework achieves state-of-the-art visual quality, reduces the number of Gaussians by eightfold, and accelerates inference speed by thirteenfold on an A100 GPU, significantly enhancing deployment efficiency on mobile devices.
📝 Abstract
We present GHARP (Real-time Gaussian Head Animation from Large-scale Reconstruction Prior), a method that animates 3D human heads in real time from a few input images of a subject and a driving expression signal. We decouple the problem into an identity stage that builds a representation of the subject's geometry and appearance offline, and an animation stage that predicts expression-dependent residuals on top of it at runtime. This separation offers a favorable trade-off with respect to fidelity, quality and runtime: the identity stage can be expensive while the animation stage runs a lightweight network, optimized for mobile devices. Our method performs animation in a semantically structured latent space of a pretrained reconstruction model, where expression changes remain spatially contained, making residual prediction efficient. This reconstruction prior provides a consistent spatial layout, allowing fusion of multiple input views into a compact, fixed-size canonical Gaussian representation. While this two-stage design improves the runtime-quality trade-off, it still inherits a problem common to all expression-driven avatar methods: expression codes describe only the face and thus omit body pose and clothing position, making these regions underspecified in the input. The animation network faces an ill-posed mapping and resorts to averaging over conflicting body appearances, producing blur and temporal flicker. We address this with a body alignment network that learns to align the person's body in the target image with the input reference images, removing the ambiguity from the training signal. Our method achieves state-of-the-art quality on the Ava-256 benchmark while running up to 13x faster on an A100 GPU with 8x fewer Gaussians.
Problem

Research questions and friction points this paper is trying to address.

3D head animation
real-time rendering
expression-driven avatar
body alignment
temporal flicker
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gaussian Head Animation
Reconstruction Prior
Latent Space Residual Prediction
Body Alignment Network
Real-time Rendering
A
Ali Benlalah
Apple
S
Sepehr Johari
Apple
P
Patricia Vitoria
Apple
Armin Kappeler
Armin Kappeler
Apple
Artem Sevastopolsky
Artem Sevastopolsky
Apple
Computer VisionMachine LearningImage ProcessingData ScienceComputer Science
A
Alexander Jung
G
Gabriele Fanelli
Kevin Mader
Kevin Mader
M
Manuel Breitenstein
C
Claudia Plüss
J
Jan Rüegg
S
Simon Biland
T
Thomas Etterlin
D
Dmitry Kostiaev
M
Mathias Deschler
B
Brian Amberg
S
Sebastian Martin