🤖 AI Summary
This study addresses the challenges of insufficient lip synchronization and audio fidelity in synchronized audio-visual generation from text or image inputs. To this end, it proposes CrossDiT, a dual-stream diffusion architecture that leverages bidirectional cross-attention to achieve temporal-semantic alignment. By incorporating a continuous pre-training strategy, the method enables synergistic cross-modal generation while preserving unimodal quality. Furthermore, supervised fine-tuning, reinforcement learning-based post-training, and model distillation techniques are introduced to support full-HD multi-modal outputs. Human evaluations demonstrate that the proposed approach yields speech quality significantly superior to previous generations and comparable to state-of-the-art models. The source code and model weights have been made publicly available.
📝 Abstract
We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (I2AV) modes; a built-in super-resolution model raises the output resolution to Full-HD (1920$\times$1080). Building on the video generation capabilities of Kandinsky 5.0, Kandinsky 6.0 Video employs a dual-stream CrossDiT architecture that connects a pretrained video stream and a newly trained audio stream through bidirectional cross-attention for temporal and semantic alignment. Our continuous pretraining strategy first trains the audio stream from scratch on large-scale audio corpora and then trains both streams jointly on paired audio-video data while preserving unimodal fidelity; pretraining is followed by supervised fine-tuning, reinforcement-learning-based post-training, and distillation. In side-by-side human evaluation, Kandinsky 6.0 Video Pro clearly outperforms its predecessor, Kandinsky 5.0 Video Pro, and remains competitive with leading audio-video generation models, particularly in speech quality. To accelerate open research and deployment in multimedia generation, we release the code, model checkpoints, and diffusers integration under the MIT license.