PixReenact: Pixel-Conditioned Causal Video Diffusion for Streaming Head-Avatar Reenactment

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of streaming head reenactment, including reliance on specialized representations, high latency from offline generation, and long-term identity drift. To overcome these challenges, this work proposes a pixel-conditioned causal video diffusion framework that directly utilizes VAE-encoded frames as driving conditions. The method incorporates cross-identity pseudo-supervision, a self-rolling anti-drift mechanism, and state-aware dual-teacher distillation to achieve low-latency real-time generation while preserving long-sequence identity consistency. Experimental results demonstrate that the proposed model exhibits strong robustness under extreme viewpoints, maintains stable identity over extended sequences, and achieves an average time-to-first-frame latency of only 239 milliseconds.
📝 Abstract
Streaming head-avatar reenactment aims to animate a reference image according to a live driving video, requiring robust motion transfer, long-term identity stability, and low latency. Existing methods often rely on specialized identity or motion representations, which can discard useful visual information and inherit failure modes from external extractors. In addition, many recent diffusion-based reenactment methods use offline, clip-based generation, jointly processing and denoising an entire video clip before producing its output, making continuous low-latency streaming difficult. We introduce PixReenact, a pixel-conditioned streaming reenactment framework built on causal video diffusion. PixReenact conditions directly on VAE-encoded reference and driving frames, without specialized identity or motion representations. To separate reference identity from driver motion, we train with cross-identity pseudo supervision together with corrective objectives anchored to the original reference and driving inputs. Long self-rollouts reduce autoregressive drift, while state-aware dual-teacher distillation separately addresses cold-start and steady-state generation. Across three cross-identity benchmarks and a long-horizon streaming benchmark, PixReenact demonstrates robust cross-identity reenactment, particularly under challenging conditions such as extreme viewpoints, occlusions, and pronounced facial expressions, while maintaining the reference identity over long streams. A 4-NFE rolling student continuously emits four frames per update with a mean emission latency of 239 ms.
Problem

Research questions and friction points this paper is trying to address.

streaming head-avatar reenactment
motion transfer
identity stability
low latency
video diffusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Causal Video Diffusion
Pixel-Conditioned Reenactment
Dual-Teacher Distillation
Streaming Head-Avatar
Cross-Identity Pseudo Supervision