🤖 AI Summary
This work addresses the limitations of diffusion-based lip synchronization models, specifically their degraded quality in complex scenes and high inference latency. To overcome these challenges, we propose a unified framework for high-fidelity, real-time generation. Methodologically, a dual-stream joint training mechanism is designed to prevent information leakage, while knowledge distillation is incorporated to accelerate single-step denoising. Furthermore, priors from vision foundation models and a relation alignment loss are introduced to enhance robustness. Experimental results demonstrate that the proposed method achieves an inference speed exceeding 70 FPS and attains state-of-the-art performance across both standard and complex scenarios. Additionally, a dedicated benchmark is constructed to facilitate evaluation. This study effectively advances the practical deployment of real-time lip synchronization technology.
📝 Abstract
Lip synchronization aims to generate visual lip dynamics that align precisely with speech audio. Despite the high generation quality of diffusion models, they often struggle in complex scenarios and suffer from prohibitive inference latency, limiting real-world deployment. We present ComplexSync, a unified diffusion-based framework that enables real-time, high-fidelity lip sync under complex conditions. First, we introduce a dual-stream joint training strategy to mitigate information leakage from reference frames while preserving natural dynamics. Second, we develop a distillation-based acceleration scheme for single-step denoising, achieving a throughput of over 70 FPS. Third, we propose a relational alignment loss that leverages structural priors from Vision Foundation Models (VFMs) to enhance robustness against complex scene factors. Furthermore, we present the first benchmark specifically designed for complex lip synchronization, comprising over 200 challenging video sequences and specialized metrics. Extensive experiments demonstrate that ComplexSync achieves state-of-the-art performance across both standard and complex scenarios while enabling real-time inference.