OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

📅 2026-07-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of fine-grained cross-modal alignment in audio-visual joint generation, which arises from structural discrepancies between modalities. To tackle this issue, the authors propose OmniVAE, a novel framework that achieves fine-grained semantic alignment in the latent space of audio and video within a variational autoencoder (VAE). OmniVAE jointly trains modality-specific VAEs while incorporating clip-level contrastive learning and knowledge distillation from pretrained modality-specific semantic encoders to enhance cross-modal consistency. Experimental results demonstrate that OmniVAE significantly improves both the quality of text-to-audio-visual generation and the accuracy of cross-modal synchronization in downstream tasks, thereby validating the critical role of unified semantic representations in holistic multimodal generative modeling.
📝 Abstract
Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to their fundamental structural differences. Most existing methods use audio and video VAEs trained separately. As a result, the two latent spaces lack cross-modal alignment, leaving the downstream generative model to learn cross-modal synchronization from scratch. We present OmniVAE, a jointly trained audio-video VAE that learns fine-grained semantic alignment between audio and video latent representations. Beyond reconstruction, OmniVAE uses a segment-level audio-video contrastive objective to capture temporal-semantic correspondence and align the two latent spaces. In parallel, it distills features from pretrained modality-specific semantic encoders into each modality, improving the downstream learnability of both latent spaces. Extensive experiments show that both objectives consistently improve the learnability of the latent spaces, translating into higher generation quality and more accurate cross-modal synchronization in downstream text-to-audio-video generation. These findings underscore the importance of learning unified representations as a foundation for omnimodal modeling.1
Problem

Research questions and friction points this paper is trying to address.

audio-video generation
cross-modal alignment
latent space
multimodal synchronization
joint generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-modal alignment
joint audio-video generation
contrastive learning
latent space distillation
omnimodal VAE
J
Jun Zhan
Fudan University; Shanghai Innovation Institute; MOSI Intelligence
Chen Yang
Chen Yang
Shanghai Jiao Tong University
Speech and Language Processing
Y
Yitian Gong
Fudan University; MOSI Intelligence
D
Donghua Yu
Fudan University; MOSI Intelligence
K
Kuangwei Chen
Fudan University; MOSI Intelligence
W
Wenbo Zhang
Fudan University; MOSI Intelligence
Kexin Huang
Kexin Huang
Fudan University
LLMAlignmentNLP
Q
Qi Luo
Fudan University; MOSI Intelligence
Z
Zhe Xu
Fudan University; MOSI Intelligence
Y
Ying Zhu
MOSI Intelligence; Shanghai Jiao Tong University
J
Jin Wang
Fudan University; MOSI Intelligence
T
Tengyue Zhang
Shanghai Innovation Institute; Shanghai Jiao Tong University
Qi Chen
Qi Chen
Xi'an Jiaotong-Liverpool University
deep learningnatural language processingsocial media data miningmachine learning applications
C
Cheng Chang
Fudan University; MOSI Intelligence
Songlin Wang
Songlin Wang
R&D Engineer, JD.com
Information RetrievalNatural Language Processing
J
Junqi Dai
Fudan University
Jiasheng Ye
Jiasheng Ye
Fudan University
Large Language ModelsGenerative ModelsAI Scientists
X
Xiaogui Yang
MOSI Intelligence
Tianyi Liang
Tianyi Liang
PHD, East China Normal University, Shanghai AI Lab,Shanghai Innovation Institute
Multimodal LearningLLMsImage Editing
X
Xiangyu Peng
MOSI Intelligence
Zhaoye Fei
Zhaoye Fei
Fudan University
Natural Language Processing
Shimin Li
Shimin Li
Fudan University
Large Language ModelSpeech Language Model
Q
Qinyuan Cheng
Fudan University; MOSI Intelligence
Xie Chen
Xie Chen
Shanghai Jiao Tong University <- Microsoft <- Cambridge University
Machine LearningSpeech RecognitionSpeech SynthesisSpeech&Audio Processing
Xinchi Chen
Xinchi Chen
Professor at Fudan University, Shanghai, China
Large Language ModelsEmbodied AINatural Language ProcessingInformation Retrievaletc.