Score
Designs and implements contrastive self-supervised pretraining pipelines that consume unlabeled raw waveforms and train encoders to produce noise‑invariant, discriminative representations optimized for label‑efficient downstream fine‑tuning and robustness on long audio segments. This includes constructing contrastive loss functions, waveform augmentations, sampling/batching strategies, and evaluation protocols for representation quality and robustness.
This work addresses the challenge of robust audio fingerprinting and retrieval for ultra-short clips (3 seconds). We propose a self-supervised contrastive learning framework based on the Conformer architecture. The method employs multi-scale feature modeling and time-aware data augmentation to learn embedding representations invariant to temporal shifts, additive noise, reverberation, and extreme time-stretching distortions. Our key contribution is a novel temporal robustness constraint mechanism that significantly enhances cross-scenario generalization. Evaluated on multiple standard audio retrieval benchmarks, the approach achieves state-of-the-art performance. Crucially, it operates directly on 3-second segments, yielding compact, efficient, and fully reproducible embeddings. All code and pre-trained models are fully open-sourced.
This work addresses the limitations of insufficient augmentation strategies and high memory consumption in audio self-supervised contrastive learning by proposing AudioMosaic, a novel contrastive learning framework based on structured time–frequency masking. By applying structured masks to spectrogram patches to construct high-quality positive pairs, AudioMosaic substantially reduces training memory usage while enhancing the discriminability and cross-domain transferability of learned representations. Integrated with large-batch training and evaluated through both linear probing and fine-tuning protocols, the method achieves state-of-the-art performance across multiple standard audio benchmarks and effectively boosts the performance of audio–language models in cross-modal tasks.
To address weak generalization and overfitting in audio few-shot and cross-domain fine-tuning—caused by the tight coupling of representation learning and classifier optimization—this paper proposes AudioConFit, the first framework to systematically integrate contrastive learning into the audio fine-tuning pipeline, thereby decoupling representation learning from classifier adaptation. Methodologically, AudioConFit combines contrastive feature alignment, a momentum encoder, temperature-scaled contrastive loss, and lightweight adapter-based fine-tuning. Evaluated on multiple standard audio classification benchmarks, it achieves state-of-the-art performance, with an average 3.2% improvement in cross-domain accuracy and a 67% reduction in trainable parameters. Crucially, it significantly enhances out-of-distribution robustness and generalization. The core contribution lies in establishing a contrastive-driven fine-tuning paradigm specifically designed for audio, offering a novel and efficient pathway for transferring pre-trained audio models to downstream tasks.
Audio representation learning heavily relies on large-scale real-world recordings, while manual annotation and data augmentation struggle to capture the full diversity of physical acoustics. Method: We propose a synthetic-driven contrastive learning framework that requires no real audio data. It employs differentiable and stochastic sound synthesizers to generate physically consistent synthetic “twin” positive pairs online via causally interpretable parameter perturbations (e.g., timbre, pitch, envelope), thereby constructing high-diversity contrastive tasks. Contribution/Results: We introduce the first positive-pair construction paradigm grounded in causal perturbations of synthesizer parameters; require only a single interpretable hyperparameter and zero real-data storage; and achieve, for the first time, synthetic-data-only models that surpass real-data baselines on ESC-50, UrbanSound8K, and SpeechCommands. This significantly reduces data dependency and storage overhead, establishing a new paradigm for low-resource audio representation learning.
Conventional self-supervised audio representation learning relies solely on clip-level sampling, leading to insufficient frame-level modeling capability. Method: This paper proposes a multi-granularity contrastive learning framework that jointly leverages clip-level, frame-level, and task-guided sampling to construct multi-perspective contrastive losses, enabling collaborative optimization of general-purpose audio representations. Contribution/Results: To our knowledge, this is the first work to incorporate both frame-level and task-specific sampling into self-supervised pre-training, overcoming the limitations of single-granularity representation learning. Pre-trained on a subset of AudioSet and evaluated via frozen-feature transfer to downstream tasks, our method achieves 25%, 20%, and 3.6% absolute improvements in clip classification, sound event detection, and pitch detection, respectively—demonstrating significantly enhanced fine-grained frame-level perception.
Existing self-supervised music foundation models exhibit limited performance on pitch-sensitive key detection tasks. This work systematically investigates how pretraining design influences pitch sensitivity and proposes a mask-contrastive pretraining approach based on Mel-spectrograms, followed by linear evaluation using a shallow, wide MLP on the learned representations. The method demonstrates for the first time that mask-contrastive embeddings substantially enhance key detection accuracy, achieving state-of-the-art results without relying on complex data augmentation. Moreover, the learned representations inherently exhibit robustness to common audio transformations, thereby validating the effectiveness of self-supervised pretraining for pitch-sensitive music information retrieval tasks.
This work addresses the representational disconnect in existing audio autoencoders between waveform reconstruction and semantic understanding, which hinders their ability to jointly excel at generation and comprehension tasks. The authors propose a unified audio tokenizer that transforms continuous latent variables into structured, generative representations through a noise-regularized bottleneck, channel normalization, and stochastic perturbation—without requiring variational training. By integrating RQ-MTP (Residual Quantization with Masked Token Prediction) training and leveraging semantic supervision from a frozen large language model, the method simultaneously optimizes high-dimensional understanding representations and continuous generative objectives. This approach achieves both high-fidelity audio reconstruction and strong semantic interpretability, effectively unifying high-quality generation with robust comprehension capabilities.
This work addresses the challenge that existing neural audio codecs struggle to balance speech intelligibility with Mel-spectrogram reconstruction fidelity, while semantic distillation approaches fail to ensure content preservation. To overcome this, we propose a self-supervised representation reconstruction (SSRR) loss, which is introduced for the first time into codec training by directly reconstructing self-supervised representations from decoded speech. This approach significantly enhances intelligibility and accelerates convergence. Integrated with a zero-lookahead streaming Transformer architecture, our method enables low-latency real-time deployment. The resulting codec, JHCodec, achieves state-of-the-art performance in both speech intelligibility and overall quality under single-GPU training conditions, and we publicly release the complete implementation and training pipeline.
This work addresses the challenge of sound event detection under limited labeled data by proposing a semi-supervised fine-tuning framework that effectively leverages abundant unlabeled data. Building upon a pretrained audio foundation model, the approach integrates pseudo-labeling, a novel conditional mixing strategy that unifies mixup and perturbation-based augmentation, and embedding-level contrastive learning. The conditional mixing mechanism harmonizes the divergent data augmentation requirements of pseudo-label learning and contrastive learning. Evaluated on the DESED validation set, the method achieves state-of-the-art performance with PSDS1 and PSDS2 scores of 0.645 and 0.822, respectively, setting a new benchmark for sound event detection in low-resource settings.
This work addresses the trade-off between event understanding and generalization capability in existing mask prediction–based audio self-supervised learning methods, which often incur high computational costs. To reconcile efficiency and effectiveness, the authors propose a lightweight Dispersion-Weighted Masking (DWM) strategy that leverages the spectral sparsity of audio spectrograms to dynamically adjust the masking distribution, thereby enhancing representation quality. By prioritizing informative yet sparse regions in the time–frequency domain, DWM significantly reduces computational complexity while consistently improving performance across multiple audio event understanding benchmarks. The approach effectively mitigates the longstanding tension between model efficiency and representational power in audio self-supervised learning.