Score
Designs, builds, or evaluates systems and algorithms that temporally align and correlate multiple signal streams—most commonly audio and video—by extracting modality-specific features (e.g., audio waveforms, visual lip motion or other signals/信号提取) and estimating precise audio-visual synchronization. This includes implementing noise-robust training and multi-signal correlation methods, developing alignment and lip-sync assessment/evaluation metrics, and producing video-processing components that synchronize audiovisual data across channels.
This paper systematically analyzes four core challenges in multimodal alignment and fusion—cross-modal misalignment, semantic modality gap, computational bottlenecks, and data noise/heterogeneity—that hinder model robustness, generalizability, and scalability. To address them, the authors propose a unified taxonomy covering 200+ studies, revealing an integrated “alignment–fusion” co-design paradigm. They further introduce three principled solution pathways: (1) noise-robust learning, (2) heterogeneous representation modeling, and (3) few-shot cross-modal transfer—unifying contrastive learning, cross-modal attention, latent-space alignment, graph neural networks, and differentiable architecture search. Extensive experiments on social media analysis, medical imaging, and sentiment recognition validate the efficacy of the framework. The work distills actionable design principles for scalable, robust, and generalizable multimodal learning, establishing a new benchmark for both theoretical research and industrial deployment.
Existing systems struggle to support strict audio-visual synchronization, limiting the analysis of fine-grained temporal features in dialogue such as turn-taking, overlapping speech, and prosody. To address this challenge, this work proposes an end-to-end multimodal acquisition and calibration framework that treats synchronized audio and video as equally central modalities for the first time. By integrating a multi-camera array with multi-channel microphones under a unified temporal architecture, the system enables scalable, reproducible, high-quality recording. Standardized calibration and quality control procedures ensure high temporal consistency across modalities, yielding data that effectively supports fine-grained analysis of conversational behavior and data-driven modeling.
This work addresses the limitations of existing audio-visual synchronization evaluation methods, which struggle to disentangle temporal alignment from semantic consistency and suffer from coupling biases in data construction. We propose the first structured and scalable benchmark framework that enables independent assessment of temporal synchronization and semantic correspondence. Through a hybrid pipeline combining automated filtering and human verification, we construct a large-scale dataset comprising 3,269 videos and 38,390 samples across three audio categories—speech, music, and environmental sounds—and ten diverse scenarios. The dataset ensures authentic on-screen sound sources and supports both multimodal alignment analysis and downstream task evaluation. Using this benchmark, we systematically evaluate five representative models. Both code and data are publicly released.
To address poor audio quality, weak semantic alignment, and audio-visual desynchronization in video-to-audio generation, this paper proposes MMAudio, a multimodal joint-training framework. MMAudio is the first to unify video-audio and text-audio dual-path generation within a single architecture. It introduces a frame-level conditional synchronization module to achieve fine-grained alignment between video features and the audio latent space, and employs flow matching as the end-to-end optimization objective. The method supports either video-only or video-plus-text conditional inputs. On public benchmarks, MMAudio achieves state-of-the-art performance: significantly improved audio fidelity, enhanced semantic alignment, and reduced audio-visual synchronization error. At inference, it generates 8-second audio clips in 1.23 seconds, with a compact model size of only 157 million parameters.
Existing audio-driven portrait animation methods rely on visual priors or local audio modeling, resulting in low motion naturalness and temporal inconsistency. This paper proposes a purely audio-driven paradigm that eliminates all visual guidance, using only raw audio as input. Our method introduces three key innovations: (1) a novel decoupled intra- and inter-utterance audio perception mechanism; (2) a context-enhanced audio encoder, a motion-decoupled controller, and a time-aware positional offset fusion module; and (3) an end-to-end framework integrating long-horizon modeling, audio-motion decoupled control, and sliding temporal window fusion. Quantitative and qualitative evaluations demonstrate state-of-the-art performance across four critical dimensions: lip-sync accuracy, video fidelity, temporal coherence, and motion diversity.
Heterogeneous camera systems—e.g., visible-light/infrared, professional/consumer-grade, or audio-equipped/audio-less setups—lack hardware synchronization in real-world scenarios, leading to significant spatiotemporal misalignment across multi-view videos. Method: This paper proposes a vision-based time-encoding method leveraging a custom-designed LED Clock. By embedding temporal exposure timestamps within frames using red and infrared LEDs, the approach achieves cross-modal, audio-free, and external-timecode-free millisecond-level synchronization. It further integrates RMSE-optimized temporal alignment with joint multi-device calibration. Contribution/Results: The method reduces synchronization residuals to 1.34 ms—substantially outperforming existing optical signal, audio-based, and timecode synchronization schemes. Validated in large-scale surgical recordings involving over 25 heterogeneous cameras, it significantly improves downstream tasks including multi-view 3D reconstruction and pose estimation.
This work addresses the challenge that certain modalities may introduce interference under specific inputs in multimodal fusion, thereby degrading model performance. To mitigate this issue, the authors propose a plug-and-play pre-fusion calibration module that leverages cross-modal summary contrast to extract supportive and conflicting cues, generating instance-level and dimension-level modulation signals. These signals dynamically enhance beneficial features while suppressing misleading information. The method achieves, for the first time, fine-grained, conflict-aware modulation prior to fusion and is compatible with both sequential and convolutional architectures. It consistently improves performance across five benchmark tasks—including emotion understanding and action recognition—and demonstrates enhanced robustness and consistency under modality missingness and data perturbations.
This work addresses the "Clever Hans effect" in existing audio-visual multimodal large language models, which often hallucinate audio based on visual cues rather than genuinely comprehending auditory content. To systematically probe whether models achieve authentic audio-visual alignment, the authors propose the Thud framework, introducing three counterfactual audio-editing interventions—Shift, Mute, and Swap—for the first time. They further design a two-stage alignment training strategy that integrates preference pairs derived from these interventions with event-level video preference regularization to enhance audio grounding. Evaluated on a 10K-sample training set, the model demonstrates a 28-percentage-point average accuracy improvement across the three intervention dimensions and achieves consistent, albeit modest, performance gains on standard video and audio-visual question-answering benchmarks.
Existing audio-visual joint generation methods struggle to simultaneously achieve fine-grained audio-visual co-evolution and tight coupling between semantic coherence and low-level synchronization. To address this challenge, this work proposes the NAVA framework, which leverages a context-conditioned native audio-visual alignment mechanism to establish cross-modal correspondences within a dedicated interaction space and guide the joint denoising process. Key innovations include an Align-then-Fuse MMDiT architecture that enables a smooth transition from modality-aware alignment to shared denoising, and a Timbre-in-Context Conditioning mechanism that supports controllable voice timbre generation. Experimental results demonstrate that, with only 6.3B parameters, NAVA significantly improves video quality, audio-visual synchronization accuracy, audio fidelity, and timbre controllability on Verse-Bench and Seed-TTS benchmarks.
Current evaluation methods for audio-visual talking head generation rely on frame-level metrics that assume strict temporal alignment between generated and reference videos, rendering them sensitive to natural variations in speech rate, rhythm, and stylistic expression, and thereby introducing assessment bias. This work reframes evaluation as a sequence alignment problem and introduces Soft Dynamic Time Warping (Soft DTW) to align feature trajectories temporally, enhancing robustness to timing offsets while preserving sequential constraints. The proposed unified sequence-level evaluation framework subsumes frame-level metrics as a special case of rigid alignment, enabling compatibility with existing perceptual, identity, and synchronization encoders without modification. Large-scale experiments across 20 methods and 7 datasets demonstrate that the approach yields more stable evaluations with higher cross-dataset consistency, clearly disentangling trade-offs between synchronization and realism, as well as expressiveness and stability.