No Distillation Needed: Single-Pass Real-Time Talking Heads via Acausal Noise Shaping

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high computational cost and streaming limitations of diffusion models in real-time audio-driven facial animation by proposing FaceGAN. This work first identifies the spectral suppression dilemma inherent in causal GANs and accordingly designs a non-causal noise shaping mechanism that leverages the future sampling properties of synthetic noise to overcome single-pass generation constraints while preserving causality along the audio pathway. Architecturally, FaceGAN employs a single-pass feedforward GAN with bounded attention windows, eliminating the need for distillation or iterative denoising. Ultimately, it achieves low-latency, infinite-length, drift-free talking-head animation—generating both expressions and head poses with a single network evaluation per frame—while matching or surpassing existing state-of-the-art methods in visual quality.
📝 Abstract
Audio-driven facial animation underpins real-time avatars, telepresence, and embodied virtual agents. And it must run online: each frame emitted from audio observed up to the current time, at interactive rates. Recent progress is dominated by diffusion models, which need many network evaluations per sample and are therefore a poor fit for streaming. We argue the cost is unnecessary in this domain. Audio-conditioned facial motion occupies a comparatively low-dimensional manifold, a regime where a single-pass GAN suffices. The obstacle is not capacity but stochastic structure. We show that a causal, time-invariant generator driven by i.i.d. noise cannot suppress its output spectrum over a band without collapsing its per-step innovation. We proposed FaceGAN, which dissolved the limitation by shaping the noise pathway acausally. Because the driving noise is synthetic, its future can be sampled now, so the audio-to-expression path stays causal, and the model supports fully causal operation. FaceGAN emits expression and head pose in a single forward pass per frame and matches or outperforms state-of-art approaches in generation quality. Being feed-forward with bounded attention windows, it generates indefinitely without drift.
Problem

Research questions and friction points this paper is trying to address.

audio-driven facial animation
real-time generation
single-pass GAN
causal generation
noise collapse
Innovation

Methods, ideas, or system contributions that make the work stand out.

Acausal Noise Shaping
Single-Pass Generation
Talking Heads
Causal Generator
FaceGAN