Score
Design and train models and pretraining pipelines that learn unified embedding spaces by aligning audio signals with textual captions through caption-based or cross-modal pretraining on audio–caption pairs; this covers methods for producing universal audio embeddings, transferring audio embeddings across domains, and extracting multimodal embeddings for retrieval, prompting, or other downstream tasks. Build, evaluate, or analyze the representation learning objectives, alignment mechanisms, and transfer techniques that ground audio semantics in text-aligned embedding spaces.
A systematic survey of Audio-Language Models (ALMs) for general-purpose audio tasks is currently lacking. Method: This paper introduces, for the first time, a six-dimensional comprehensive taxonomy—covering architectural design, pretraining paradigms, downstream adaptation strategies, benchmark datasets, evaluation protocols, and future challenges—and proposes a structured technical roadmap. Our methodology integrates multimodal representation learning, contrastive and generative pretraining, instruction tuning, multi-task collaborative optimization, and agent-based system design. Contribution/Results: The work fills a critical gap in the ALM literature by delivering the first authoritative, holistic survey; it provides researchers and practitioners with a rigorous technical reference and practical guidance, thereby significantly advancing human-like auditory modeling research and real-world applications.
Audio-language pretraining suffers from insufficient large-scale high-quality data, limited caption diversity, and a lack of systematic evaluation—hindering progress toward general-purpose audio understanding. To address these challenges, we introduce CaptionStew, a multi-style audio-text dataset, and conduct the first systematic comparative study of contrastive versus generative caption modeling for cross-domain representation learning across speech, music, and environmental sounds. Leveraging data augmentation and large-scale joint pretraining, we develop a unified audio encoder. Experiments demonstrate substantial gains in transfer performance on diverse downstream tasks, with particularly pronounced improvements in low-resource settings. We fully open-source our data curation pipelines, training code, and pretrained models—establishing critical infrastructure and an empirical benchmark for audio-language foundation model research.
Existing two-stage cross-modal alignment methods suffer from suboptimal semantic alignment due to distributional mismatches across modalities. To address this, we propose the first single-stage, trimodal joint contrastive learning framework for end-to-end semantic alignment of audio, visual, and textual modalities. Our approach abandons the sequential alignment paradigm, instead constructing a unified representation space via cross-modal attention and multimodal embedding. A novel triplet loss is introduced to enhance contrastive learning across all three modalities simultaneously. Evaluated on the AVCaps dataset, our method achieves the first empirical validation of single-stage alignment superiority: audio-driven visual retrieval improves by 2× over two-stage baselines, and cross-modal retrieval consistently outperforms state-of-the-art two-stage models across all modalities. These results demonstrate both the effectiveness and scalability of unified multimodal representation learning.
Current contrastive language–audio pretraining models (e.g., CLAP) rely on global audio descriptions and lack fine-grained temporal supervision, limiting their frame-level alignment capability. To address this, we propose: (1) the first large-scale temporally aligned audio–text dataset—comprising 12k Freesound clips—with each text caption precisely annotated to a corresponding audio subsegment; (2) a temporal segment-level annotation paradigm coupled with frame-level contrastive learning, enhanced by LLM-driven annotation cleaning to ensure high quality; and (3) an extended CLAP architecture with a temporally aware contrastive loss. Our approach achieves significant improvements on the AudioSet Strong benchmark in both temporal localization and local alignment, demonstrating the critical role of strong temporal supervision in language–audio joint modeling. The dataset, code, and models are publicly released to advance fine-grained cross-modal understanding.
Existing general-purpose audio pretraining is constrained by weak, noisy, and limited-scale labels, lacking a unified strong supervision framework. This work proposes the first Unified Tag System (UTS) that integrates speech, music, and environmental sounds, and establishes a high-fidelity audio captioning pipeline to enable a new pretraining paradigm centered on high-quality, strongly supervised data. Through systematic evaluation of multiple pretraining objectives within this framework, the study demonstrates that data quality and coverage are critical to performance gains, and further reveals that different pretraining objectives substantially influence the model’s specialization capabilities across downstream tasks.
To address poor audio quality, weak semantic alignment, and audio-visual desynchronization in video-to-audio generation, this paper proposes MMAudio, a multimodal joint-training framework. MMAudio is the first to unify video-audio and text-audio dual-path generation within a single architecture. It introduces a frame-level conditional synchronization module to achieve fine-grained alignment between video features and the audio latent space, and employs flow matching as the end-to-end optimization objective. The method supports either video-only or video-plus-text conditional inputs. On public benchmarks, MMAudio achieves state-of-the-art performance: significantly improved audio fidelity, enhanced semantic alignment, and reduced audio-visual synchronization error. At inference, it generates 8-second audio clips in 1.23 seconds, with a compact model size of only 157 million parameters.
This work addresses the modality gap in audio–text multimodal contrastive embeddings, which limits zero-shot task performance. Departing from conventional views that attribute this gap primarily to mean shift, the study reveals that similarity computation is instead dominated by a few interpretable concept axes within a decomposed conceptual space. Building on this insight, the authors propose an untrained spectral truncation method that leverages partial least squares singular value decomposition (PLS-SVD) to disentangle the embedding space and identify these critical concept axes. Without requiring large memory banks or additional training, the approach substantially reduces embedding dimensionality while achieving zero-shot audio captioning performance approaching that of fully supervised methods, and maintains competitive results in both retrieval and generation tasks.
研究通过音频-描述对齐方法改进预训练编码器,提升跨域分类准确性,使用线性探针和序列感知LLM读出进行评估。
This work addresses the limitations of existing audio–text retrieval methods, which suffer from degraded performance on long-duration, noisy, and weakly labeled audio and exhibit instability under small-batch training. To overcome these challenges, the authors propose a cross-modal embedding refinement mechanism that integrates Transformer-based projection, linear mapping, and bidirectional attention. Additionally, they introduce a silence-aware chunking strategy coupled with attentive pooling to better capture relevant audio segments. A hybrid loss function combining cosine similarity, L1 regularization, and contrastive loss is designed to enhance model robustness and training stability. Experimental results demonstrate that the proposed approach significantly outperforms state-of-the-art methods on standard benchmarks, exhibiting particularly strong robustness in noisy conditions with signal-to-noise ratios between 5 and 15 dB.
This work proposes a general-purpose audio embedding framework that addresses the limitations of existing audio-text retrieval models, which are typically optimized solely for caption matching and struggle to support diverse objectives or controllable retrieval. For the first time, natural language instructions are integrated into the embedding process, leveraging a pretrained large audio-language model to transfer its capabilities in audio understanding, instruction following, and reasoning. The approach employs a contrastive dual-encoder architecture and utilizes large-scale multimodal pretraining to facilitate embedding space transfer learning. Experimental results demonstrate that the model achieves strong performance on standard audio and speech retrieval benchmarks while exhibiting remarkable compositional generalization and instruction-controllable retrieval abilities, underscoring its potential as a universal audio embedding model.
Existing multimodal embedding models struggle to uniformly support text, images, video, and audio, often omitting audio or covering only a subset of modalities. This work proposes a unified four-modality embedding space built upon a frozen vision–language foundation model, augmented with a lightweight audio tower connector and modality-gated deep adapters, requiring no updates to the base model parameters. By aligning only audio with text, the approach enables cross-modal retrieval across all pairs—including audio–image—while fully preserving the original performance on pre-existing modalities. Both generations of the proposed model are trained within hours on a single GPU and achieve strong results in audio–text and audio–image retrieval. The code, model weights, and evaluation tools are publicly released.