PixReenact: Pixel-Conditioned Causal Video Diffusion for Streaming Head-Avatar Reenactment
This study addresses the limitations of streaming head reenactment, including reliance on specialized representations, high latency from offline generation, and long-term identity drift. To overcome these challenges, this work proposes a pixel-conditioned causal video diffusion framework that directly utilizes VAE-encoded frames as driving conditions. The method incorporates cross-identity pseudo-supervision, a self-rolling anti-drift mechanism, and state-aware dual-teacher distillation to achieve low-latency real-time generation while preserving long-sequence identity consistency. Experimental results demonstrate that the proposed model exhibits strong robustness under extreme viewpoints, maintains stable identity over extended sequences, and achieves an average time-to-first-frame latency of only 239 milliseconds.