Score
Designs and implements model architectures that first compute audio–visual alignment—including temporal synchronization, correspondence, and modality-aware representations—optionally conditioned on external context, and then perform staged fusion of the modalities. These competencies are used to build or analyze multimodal pipelines (for example joint denoising or inference) in which synchronization and alignment are decoupled from semantic conditioning and fusion occurs as a separate, later stage.
This paper systematically analyzes four core challenges in multimodal alignment and fusion—cross-modal misalignment, semantic modality gap, computational bottlenecks, and data noise/heterogeneity—that hinder model robustness, generalizability, and scalability. To address them, the authors propose a unified taxonomy covering 200+ studies, revealing an integrated “alignment–fusion” co-design paradigm. They further introduce three principled solution pathways: (1) noise-robust learning, (2) heterogeneous representation modeling, and (3) few-shot cross-modal transfer—unifying contrastive learning, cross-modal attention, latent-space alignment, graph neural networks, and differentiable architecture search. Extensive experiments on social media analysis, medical imaging, and sentiment recognition validate the efficacy of the framework. The work distills actionable design principles for scalable, robust, and generalizable multimodal learning, establishing a new benchmark for both theoretical research and industrial deployment.
Existing research lacks a systematic analysis of modality–language backbone integration mechanisms in multimodal large language models (MLLMs). Method: We systematically survey 125 MLLMs published between 2021 and 2025, proposing the first large-language-model-centric three-dimensional taxonomy—architectural integration, representation learning, and training paradigms—to unify cross-modal alignment pathways. Leveraging bibliometric analysis, architectural decoupling, and fine-grained modeling of embeddings and loss functions, we characterize evolutionary patterns in fusion granularity, joint representation design, and objective function development. Contribution/Results: This work fills a critical theoretical gap in structured analysis of MLLM fusion mechanisms, identifies shared bottlenecks—including misaligned modality-specific feature hierarchies and suboptimal cross-modal optimization objectives—and delivers a reusable theoretical framework and practical guidelines for developing robust, scalable multimodal foundation models.
Existing two-stage cross-modal alignment methods suffer from suboptimal semantic alignment due to distributional mismatches across modalities. To address this, we propose the first single-stage, trimodal joint contrastive learning framework for end-to-end semantic alignment of audio, visual, and textual modalities. Our approach abandons the sequential alignment paradigm, instead constructing a unified representation space via cross-modal attention and multimodal embedding. A novel triplet loss is introduced to enhance contrastive learning across all three modalities simultaneously. Evaluated on the AVCaps dataset, our method achieves the first empirical validation of single-stage alignment superiority: audio-driven visual retrieval improves by 2× over two-stage baselines, and cross-modal retrieval consistently outperforms state-of-the-art two-stage models across all modalities. These results demonstrate both the effectiveness and scalability of unified multimodal representation learning.
To address poor audio quality, weak semantic alignment, and audio-visual desynchronization in video-to-audio generation, this paper proposes MMAudio, a multimodal joint-training framework. MMAudio is the first to unify video-audio and text-audio dual-path generation within a single architecture. It introduces a frame-level conditional synchronization module to achieve fine-grained alignment between video features and the audio latent space, and employs flow matching as the end-to-end optimization objective. The method supports either video-only or video-plus-text conditional inputs. On public benchmarks, MMAudio achieves state-of-the-art performance: significantly improved audio fidelity, enhanced semantic alignment, and reduced audio-visual synchronization error. At inference, it generates 8-second audio clips in 1.23 seconds, with a compact model size of only 157 million parameters.
The core challenge in multimodal representation learning lies in modeling comparability across inherently incomparable modalities. This paper investigates why large-scale unimodal models spontaneously develop cross-modal representation alignment—even without explicit alignment supervision—and systematically examines the necessity and efficacy of such alignment. We propose an empirical framework integrating representation similarity measurement, information decomposition analysis, and performance attribution. Our results—first to rigorously test alignment’s universal benefit—show that alignment is not inherently advantageous: its utility critically depends on the dynamic interplay between inter-modal semantic similarity and the balance of redundant versus unique information. Under certain data conditions, excessive alignment degrades downstream performance. The study identifies the emergent conditions for implicit alignment and clarifies its true relationship with task performance, challenging the prevailing “alignment-is-better” assumption. These findings provide both theoretical grounding and practical guidelines for data-driven alignment strategy design.
Open-source audio-visual generation models suffer from unstable cross-modal synchronization, primarily due to three deficiencies in diffusion-based generation: (1) cross-modal correspondence drift; (2) global attention’s inability to capture fine-grained temporal alignment; and (3) intra-modal bias in conventional Classifier-Free Guidance (CFG), which undermines cross-modal synergy. To address these issues, we propose a unified framework comprising: (i) a cross-task co-training paradigm that jointly optimizes bidirectional audio-to-video and video-to-audio generation; (ii) a global-local decoupled interaction module that separately models long-range dependencies and frame-level synchronization; and (iii) a synchronization-enhanced CFG incorporating cross-modal consistency constraints to mitigate modality-specific biases. Our method achieves significant improvements in fine-grained audio-visual synchronization accuracy and generation fidelity across multiple benchmarks, establishing new state-of-the-art performance.
Existing audio-visual large language models (AV-LLMs) suffer from weak audio understanding, leading to modality hallucination and cross-modal inconsistency. To address this, we propose Dolphin, a fine-grained audio-visual co-alignment architecture, and AVU—the first open-domain, question-answering–style audio-visual instruction dataset comprising 5.2 million samples. Methodologically, Dolphin introduces a multi-scale audio-visual adapter for spatial alignment, an interleaved audio-visual fusion mechanism for temporal alignment, and a unified video-audio-question triplet encoding framework. Extensive experiments demonstrate that Dolphin achieves state-of-the-art performance across multiple audio-visual understanding benchmarks. It significantly improves factual accuracy and cross-modal consistency while effectively mitigating modality hallucination.
This work addresses the limitations of existing audio-visual synchronization evaluation methods, which struggle to disentangle temporal alignment from semantic consistency and suffer from coupling biases in data construction. We propose the first structured and scalable benchmark framework that enables independent assessment of temporal synchronization and semantic correspondence. Through a hybrid pipeline combining automated filtering and human verification, we construct a large-scale dataset comprising 3,269 videos and 38,390 samples across three audio categories—speech, music, and environmental sounds—and ten diverse scenarios. The dataset ensures authentic on-screen sound sources and supports both multimodal alignment analysis and downstream task evaluation. Using this benchmark, we systematically evaluate five representative models. Both code and data are publicly released.
Existing audio-visual joint generation methods struggle to simultaneously achieve fine-grained audio-visual co-evolution and tight coupling between semantic coherence and low-level synchronization. To address this challenge, this work proposes the NAVA framework, which leverages a context-conditioned native audio-visual alignment mechanism to establish cross-modal correspondences within a dedicated interaction space and guide the joint denoising process. Key innovations include an Align-then-Fuse MMDiT architecture that enables a smooth transition from modality-aware alignment to shared denoising, and a Timbre-in-Context Conditioning mechanism that supports controllable voice timbre generation. Experimental results demonstrate that, with only 6.3B parameters, NAVA significantly improves video quality, audio-visual synchronization accuracy, audio fidelity, and timbre controllability on Verse-Bench and Seed-TTS benchmarks.
This work addresses the lack of a systematic understanding of when cross-modal alignment (CA) and cross-modal prediction (CP) are effective in multimodal learning—a gap that often leads to suboptimal performance or even degradation relative to unimodal baselines. The authors propose a unified linear analytical framework under a structured signal–noise model with correlated interference, revealing complementary failure mechanisms of CA and CP. They introduce the first multimodal “phase diagram,” which delineates four distinct regimes: both methods succeed, only alignment works, only prediction works, or neither is effective. Leveraging separation ratio analysis, a unidirectional whitening mechanism, and a few-shot label-guided localization algorithm, this phase diagram enables practical guidance for method selection on real-world data. Experiments across synthetic, stereo vision, image–text, and astrophysical datasets validate its efficacy in identifying harmful multimodal configurations, offering a diagnostic tool for practitioners prior to model deployment.
This work addresses the challenge of fine-grained cross-modal alignment in audio-visual joint generation, which arises from structural discrepancies between modalities. To tackle this issue, the authors propose OmniVAE, a novel framework that achieves fine-grained semantic alignment in the latent space of audio and video within a variational autoencoder (VAE). OmniVAE jointly trains modality-specific VAEs while incorporating clip-level contrastive learning and knowledge distillation from pretrained modality-specific semantic encoders to enhance cross-modal consistency. Experimental results demonstrate that OmniVAE significantly improves both the quality of text-to-audio-visual generation and the accuracy of cross-modal synchronization in downstream tasks, thereby validating the critical role of unified semantic representations in holistic multimodal generative modeling.
This work addresses the challenge that certain modalities may introduce interference under specific inputs in multimodal fusion, thereby degrading model performance. To mitigate this issue, the authors propose a plug-and-play pre-fusion calibration module that leverages cross-modal summary contrast to extract supportive and conflicting cues, generating instance-level and dimension-level modulation signals. These signals dynamically enhance beneficial features while suppressing misleading information. The method achieves, for the first time, fine-grained, conflict-aware modulation prior to fusion and is compatible with both sequential and convolutional architectures. It consistently improves performance across five benchmark tasks—including emotion understanding and action recognition—and demonstrates enhanced robustness and consistency under modality missingness and data perturbations.