Proxy Avatar Meets Low-Rank Caching: Real-Time One-Shot Emotion-Controllable Portrait Animation

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing approaches to real-time, one-shot, emotion-controllable portrait animation are hindered by insufficient emotion-aware motion priors and the high computational cost of appearance modeling in multi-step denoising. This work proposes a cascaded framework that first generates an emotion-driven video using a Gaussian-based emotional proxy avatar, then transfers the identity-agnostic motion to any target portrait via a large-scale one-shot redirection model. The method introduces three key innovations: a reusable emotional proxy avatar, a zero-shot appearance feature caching mechanism, and low-rank adapters. Together, these components significantly enhance emotional expressiveness while drastically reducing inference cost, all without compromising identity consistency. To the best of our knowledge, this is the first approach to achieve efficient, real-time generation of emotion-controllable portrait animation from a single reference image.
📝 Abstract
Audio-driven portrait animation has advanced rapidly with diffusion-based generative models, yet real-time one-shot generation with expressive emotion control remains challenging. Existing methods often suffer from insufficient emotion-aware motion priors and expensive appearance computation during multi-step denoising. To address these issues, we propose Proxy Avatar Meets Low-Rank Caching, a cascaded framework for real-time one-shot emotion-controllable portrait animation. Instead of directly generating the target portrait from audio, our method uses a Gaussian-based emotion proxy avatar as a reusable motion generator, which is trained once on a single identity to produce expressive driving videos from audio and emotion labels. Since the proxy avatar only provides motion rather than target appearance or geometry, a large-scale one-shot retargeting model further extracts identity-independent motion from the proxy performance and adapts it to arbitrary target portraits. To improve inference efficiency, we introduce zero-shot appearance reuse with low-rank caching, which caches reference appearance features at the initial denoising step and models subsequent feature variations using lightweight low-rank adapters. Extensive experiments demonstrate that our method achieves stronger emotional expressiveness, better identity-preserving animation, and substantially reduced inference cost, enabling real-time one-shot portrait animation.
Problem

Research questions and friction points this paper is trying to address.

portrait animation
emotion control
real-time generation
one-shot learning
audio-driven animation
Innovation

Methods, ideas, or system contributions that make the work stand out.

emotion-controllable animation
proxy avatar
low-rank caching
one-shot portrait animation
real-time diffusion
🔎 Similar Papers
No similar papers found.