ComplexSync: High-Fidelity and Real-Time Lip Sync in Complex Scenarios
This work addresses the limitations of diffusion-based lip synchronization models, specifically their degraded quality in complex scenes and high inference latency. To overcome these challenges, we propose a unified framework for high-fidelity, real-time generation. Methodologically, a dual-stream joint training mechanism is designed to prevent information leakage, while knowledge distillation is incorporated to accelerate single-step denoising. Furthermore, priors from vision foundation models and a relation alignment loss are introduced to enhance robustness. Experimental results demonstrate that the proposed method achieves an inference speed exceeding 70 FPS and attains state-of-the-art performance across both standard and complex scenarios. Additionally, a dedicated benchmark is constructed to facilitate evaluation. This study effectively advances the practical deployment of real-time lip synchronization technology.