Score
Design, build, and evaluate models and training objectives that align audio and visual embedding spaces using contrastive losses so that semantically matching audio and visual inputs are mapped close together while mismatched pairs are separated. This includes constructing positive and negative pairs, choosing sampling and temperature strategies, and analyzing alignment quality, modality shortcut reduction, robustness to noise or interference, and impacts on downstream reconstruction or inference performance.
This paper systematically analyzes four core challenges in multimodal alignment and fusion—cross-modal misalignment, semantic modality gap, computational bottlenecks, and data noise/heterogeneity—that hinder model robustness, generalizability, and scalability. To address them, the authors propose a unified taxonomy covering 200+ studies, revealing an integrated “alignment–fusion” co-design paradigm. They further introduce three principled solution pathways: (1) noise-robust learning, (2) heterogeneous representation modeling, and (3) few-shot cross-modal transfer—unifying contrastive learning, cross-modal attention, latent-space alignment, graph neural networks, and differentiable architecture search. Extensive experiments on social media analysis, medical imaging, and sentiment recognition validate the efficacy of the framework. The work distills actionable design principles for scalable, robust, and generalizable multimodal learning, establishing a new benchmark for both theoretical research and industrial deployment.
Existing two-stage cross-modal alignment methods suffer from suboptimal semantic alignment due to distributional mismatches across modalities. To address this, we propose the first single-stage, trimodal joint contrastive learning framework for end-to-end semantic alignment of audio, visual, and textual modalities. Our approach abandons the sequential alignment paradigm, instead constructing a unified representation space via cross-modal attention and multimodal embedding. A novel triplet loss is introduced to enhance contrastive learning across all three modalities simultaneously. Evaluated on the AVCaps dataset, our method achieves the first empirical validation of single-stage alignment superiority: audio-driven visual retrieval improves by 2× over two-stage baselines, and cross-modal retrieval consistently outperforms state-of-the-art two-stage models across all modalities. These results demonstrate both the effectiveness and scalability of unified multimodal representation learning.
This study systematically compares contrastive loss and triplet loss in audio-visual cross-modal embedding learning, focusing on their representational capacity differences. To characterize their intrinsic distinctions—particularly in intra-class variance control, hard sample mining, and optimization dynamics—we propose a quantitative analysis framework measuring loss decay rate, positive-pair activation ratio, and gradient norm. Controlled experiments across MNIST, CIFAR-10, CUB-200, CARS196, and synthetic datasets demonstrate that triplet loss prioritizes hard samples, preserves richer semantic details, and significantly improves fine-grained classification and cross-modal retrieval performance; contrastive loss yields more compact yet less discriminative embeddings. Crucially, this work provides the first optimization-dynamics–driven mechanistic explanation of how these losses differentially shape representation quality—establishing both theoretical grounding and empirical evidence for principled loss selection in metric learning.
Existing methods jointly optimize contrastive alignment and masked reconstruction objectives, which often introduces semantic noise and causes optimization interference, thereby limiting cross-modal representation learning performance. This work proposes the TG-DP framework, which decouples reconstruction and alignment tasks along separate optimization paths for the first time. Each path employs a visibility pattern tailored to its specific objective, and a teacher model is introduced to guide the organization of visible tokens in the contrastive path, reducing interference and enhancing representation quality. The proposed method achieves significant improvements in zero-shot retrieval on AudioSet—R@1 increases from 35.2% to 37.4% (video→audio) and from 27.9% to 37.1% (audio→video)—and attains state-of-the-art linear probe performance on both AS20K and VGGSound benchmarks.
This work addresses cross-modal video-to-audio generation with emphasis on semantic consistency and frame-level temporal alignment. To this end, we propose an alignment-aware framework featuring: (i) a lightweight visual encoder for efficient video representation extraction; (ii) learnable auxiliary embeddings that explicitly model audio–video correspondence; and (iii) multi-scale temporal data augmentation coupled with end-to-end joint training to enforce temporal coherence. Our key innovation lies in an implicit alignment mechanism, which reveals the critical role of auxiliary embeddings and augmentation strategies in achieving precise synchronization. We further introduce the first comprehensive evaluation paradigm specifically designed for audio–video alignment. Experiments demonstrate state-of-the-art performance in both audio fidelity—measured by STFT-L1 and PESQ—and frame-level synchronization accuracy—quantified by SyncScore—establishing a new benchmark for photorealistic audiovisual generation.
This work addresses the challenges of spurious negative samples and missing cross-modal semantic associations in audio-visual embedding learning caused by sparse annotations. To mitigate these issues, we propose a novel learning framework that leverages soft-label prediction and an implicit interaction graph. Our approach employs a teacher–student architecture to generate reliable soft supervision signals and utilizes the GRaSP algorithm to construct a directed inter-class dependency graph. By incorporating graph-guided regularization and semantic alignment losses, the model effectively captures latent semantic dependencies among unannotated co-occurring events. Experiments on the AVE and VEGAS benchmarks demonstrate that the proposed method significantly improves mean average precision (mAP), enhancing both semantic consistency and robustness in cross-modal embeddings.
This work addresses the limitations of existing audio–text retrieval methods, which suffer from degraded performance on long-duration, noisy, and weakly labeled audio and exhibit instability under small-batch training. To overcome these challenges, the authors propose a cross-modal embedding refinement mechanism that integrates Transformer-based projection, linear mapping, and bidirectional attention. Additionally, they introduce a silence-aware chunking strategy coupled with attentive pooling to better capture relevant audio segments. A hybrid loss function combining cosine similarity, L1 regularization, and contrastive loss is designed to enhance model robustness and training stability. Experimental results demonstrate that the proposed approach significantly outperforms state-of-the-art methods on standard benchmarks, exhibiting particularly strong robustness in noisy conditions with signal-to-noise ratios between 5 and 15 dB.
研究通过音频-描述对齐方法改进预训练编码器,提升跨域分类准确性,使用线性探针和序列感知LLM读出进行评估。
This work addresses the modality gap in audio–text multimodal contrastive embeddings, which limits zero-shot task performance. Departing from conventional views that attribute this gap primarily to mean shift, the study reveals that similarity computation is instead dominated by a few interpretable concept axes within a decomposed conceptual space. Building on this insight, the authors propose an untrained spectral truncation method that leverages partial least squares singular value decomposition (PLS-SVD) to disentangle the embedding space and identify these critical concept axes. Without requiring large memory banks or additional training, the approach substantially reduces embedding dimensionality while achieving zero-shot audio captioning performance approaching that of fully supervised methods, and maintains competitive results in both retrieval and generation tasks.
This work addresses the significant distributional discrepancy and structural misalignment between audio-visual and textual modalities in generalized zero-shot learning (GZSL). To mitigate these challenges, the authors propose a novel approach that integrates Z-score normalization with a three-level hierarchical alignment mechanism. By normalizing fused audio-visual and textual embeddings and jointly optimizing alignment at the semantic, class, and batch levels within a shared embedding space, the method effectively reduces inter-modal distributional shifts while preserving semantic relationships and intra-batch spatial consistency. Evaluated on three standard benchmarks—VGGSound-GZSL, UCF-GZSL, and ActivityNet-GZSL—the proposed framework achieves competitive performance. Notably, it is the first to incorporate both normalization and multi-level alignment into audio-visual GZSL, substantially enhancing the robustness and generalization of learned representations.
This work proposes a general-purpose audio embedding framework that addresses the limitations of existing audio-text retrieval models, which are typically optimized solely for caption matching and struggle to support diverse objectives or controllable retrieval. For the first time, natural language instructions are integrated into the embedding process, leveraging a pretrained large audio-language model to transfer its capabilities in audio understanding, instruction following, and reasoning. The approach employs a contrastive dual-encoder architecture and utilizes large-scale multimodal pretraining to facilitate embedding space transfer learning. Experimental results demonstrate that the model achieves strong performance on standard audio and speech retrieval benchmarks while exhibiting remarkable compositional generalization and instruction-controllable retrieval abilities, underscoring its potential as a universal audio embedding model.