contrastive waveform pretraining

Designs and implements contrastive self-supervised pretraining pipelines that consume unlabeled raw waveforms and train encoders to produce noise‑invariant, discriminative representations optimized for label‑efficient downstream fine‑tuning and robustness on long audio segments. This includes constructing contrastive loss functions, waveform augmentations, sampling/batching strategies, and evaluation protocols for representation quality and robustness.

contrastivewaveformpretraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.8
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Pretrained Conformers for Audio Fingerprinting and Retrieval

Aug 15, 2025
KA
Kemal Altwlkany
🏛️ Infobip | University of Sarajevo

This work addresses the challenge of robust audio fingerprinting and retrieval for ultra-short clips (3 seconds). We propose a self-supervised contrastive learning framework based on the Conformer architecture. The method employs multi-scale feature modeling and time-aware data augmentation to learn embedding representations invariant to temporal shifts, additive noise, reverberation, and extreme time-stretching distortions. Our key contribution is a novel temporal robustness constraint mechanism that significantly enhances cross-scenario generalization. Evaluated on multiple standard audio retrieval benchmarks, the approach achieves state-of-the-art performance. Crucially, it operates directly on 3-second segments, yielding compact, efficient, and fully reproducible embeddings. All code and pre-trained models are fully open-sourced.

Achieve robust audio retrieval with short 3-second segmentsDevelop conformer-based encoders for audio fingerprintingImprove resistance to noise, reverb, and temporal distortions

This work addresses the limitations of insufficient augmentation strategies and high memory consumption in audio self-supervised contrastive learning by proposing AudioMosaic, a novel contrastive learning framework based on structured time–frequency masking. By applying structured masks to spectrogram patches to construct high-quality positive pairs, AudioMosaic substantially reduces training memory usage while enhancing the discriminability and cross-domain transferability of learned representations. Integrated with large-batch training and evaluated through both linear probing and fine-tuning protocols, the method achieves state-of-the-art performance across multiple standard audio benchmarks and effectively boosts the performance of audio–language models in cross-modal tasks.

audio representation learningaudio self-supervised learningcontrastive learning

Audio Contrastive based Fine-tuning

Sep 21, 2023
YW
Yang Wang
🏛️ The University of Sheffield | Rainbow-Flower AI | Durham University | The University of Manchester

To address weak generalization and overfitting in audio few-shot and cross-domain fine-tuning—caused by the tight coupling of representation learning and classifier optimization—this paper proposes AudioConFit, the first framework to systematically integrate contrastive learning into the audio fine-tuning pipeline, thereby decoupling representation learning from classifier adaptation. Methodologically, AudioConFit combines contrastive feature alignment, a momentum encoder, temperature-scaled contrastive loss, and lightweight adapter-based fine-tuning. Evaluated on multiple standard audio classification benchmarks, it achieves state-of-the-art performance, with an average 3.2% improvement in cross-domain accuracy and a 67% reduction in trainable parameters. Crucially, it significantly enhances out-of-distribution robustness and generalization. The core contribution lies in establishing a contrastive-driven fine-tuning paradigm specifically designed for audio, offering a novel and efficient pathway for transferring pre-trained audio models to downstream tasks.

Decoupling representation learning from classifier training in audio modelsEvaluating representation quality using dual-probe geometric analysisImproving geometric structure of embedding space via contrastive-tuning

Contrastive Learning from Synthetic Audio Doppelgangers

Jun 09, 2024
MC
Manuel Cherep
🏛️ Massachusetts Institute of Technology

Audio representation learning heavily relies on large-scale real-world recordings, while manual annotation and data augmentation struggle to capture the full diversity of physical acoustics. Method: We propose a synthetic-driven contrastive learning framework that requires no real audio data. It employs differentiable and stochastic sound synthesizers to generate physically consistent synthetic “twin” positive pairs online via causally interpretable parameter perturbations (e.g., timbre, pitch, envelope), thereby constructing high-diversity contrastive tasks. Contribution/Results: We introduce the first positive-pair construction paradigm grounded in causal perturbations of synthesizer parameters; require only a single interpretable hyperparameter and zero real-data storage; and achieve, for the first time, synthetic-data-only models that surpass real-data baselines on ESC-50, UrbanSound8K, and SpeechCommands. This significantly reduces data dependency and storage overhead, establishing a new paradigm for low-resource audio representation learning.

Existing audio transformations lack true real-world sound diversity.Learning robust audio representations requires large real-world datasets.Synthetic audio doppelgängers provide rich contrastive learning information.

Self-Supervised Learning Method Using Multiple Sampling Strategies for General-Purpose Audio Representation

May 23, 2022
IK
Ibuki Kuroyanagi
🏛️ LINE Corporation | Nagoya University

Conventional self-supervised audio representation learning relies solely on clip-level sampling, leading to insufficient frame-level modeling capability. Method: This paper proposes a multi-granularity contrastive learning framework that jointly leverages clip-level, frame-level, and task-guided sampling to construct multi-perspective contrastive losses, enabling collaborative optimization of general-purpose audio representations. Contribution/Results: To our knowledge, this is the first work to incorporate both frame-level and task-specific sampling into self-supervised pre-training, overcoming the limitations of single-granularity representation learning. Pre-trained on a subset of AudioSet and evaluated via frozen-feature transfer to downstream tasks, our method achieves 25%, 20%, and 3.6% absolute improvements in clip classification, sound event detection, and pitch detection, respectively—demonstrating significantly enhanced fine-grained frame-level perception.

Develop self-supervised learning for general audio representationEnhance pitch detection and sound event detection performanceImprove frame-level classification via multi-strategy contrastive losses

Latest Papers

What's happening recently
View more

Existing self-supervised music foundation models exhibit limited performance on pitch-sensitive key detection tasks. This work systematically investigates how pretraining design influences pitch sensitivity and proposes a mask-contrastive pretraining approach based on Mel-spectrograms, followed by linear evaluation using a shallow, wide MLP on the learned representations. The method demonstrates for the first time that mask-contrastive embeddings substantially enhance key detection accuracy, achieving state-of-the-art results without relying on complex data augmentation. Moreover, the learned representations inherently exhibit robustness to common audio transformations, thereby validating the effectiveness of self-supervised pretraining for pitch-sensitive music information retrieval tasks.

key detectionmusic audiomusic information retrieval

This work addresses the representational disconnect in existing audio autoencoders between waveform reconstruction and semantic understanding, which hinders their ability to jointly excel at generation and comprehension tasks. The authors propose a unified audio tokenizer that transforms continuous latent variables into structured, generative representations through a noise-regularized bottleneck, channel normalization, and stochastic perturbation—without requiring variational training. By integrating RQ-MTP (Residual Quantization with Masked Token Prediction) training and leveraging semantic supervision from a frozen large language model, the method simultaneously optimizes high-dimensional understanding representations and continuous generative objectives. This approach achieves both high-fidelity audio reconstruction and strong semantic interpretability, effectively unifying high-quality generation with robust comprehension capabilities.

audio autoencoderaudio tokenizerlatent representation

This work addresses the challenge that existing neural audio codecs struggle to balance speech intelligibility with Mel-spectrogram reconstruction fidelity, while semantic distillation approaches fail to ensure content preservation. To overcome this, we propose a self-supervised representation reconstruction (SSRR) loss, which is introduced for the first time into codec training by directly reconstructing self-supervised representations from decoded speech. This approach significantly enhances intelligibility and accelerates convergence. Integrated with a zero-lookahead streaming Transformer architecture, our method enables low-latency real-time deployment. The resulting codec, JHCodec, achieves state-of-the-art performance in both speech intelligibility and overall quality under single-GPU training conditions, and we publicly release the complete implementation and training pipeline.

low-latency streamingneural audio codecrepresentation reconstruction

This work addresses the challenge of sound event detection under limited labeled data by proposing a semi-supervised fine-tuning framework that effectively leverages abundant unlabeled data. Building upon a pretrained audio foundation model, the approach integrates pseudo-labeling, a novel conditional mixing strategy that unifies mixup and perturbation-based augmentation, and embedding-level contrastive learning. The conditional mixing mechanism harmonizes the divergent data augmentation requirements of pseudo-label learning and contrastive learning. Evaluated on the DESED validation set, the method achieves state-of-the-art performance with PSDS1 and PSDS2 scores of 0.645 and 0.822, respectively, setting a new benchmark for sound event detection in low-resource settings.

Fine-tuningLabeled Data ScarcitySemi-Supervised Learning

This work addresses the trade-off between event understanding and generalization capability in existing mask prediction–based audio self-supervised learning methods, which often incur high computational costs. To reconcile efficiency and effectiveness, the authors propose a lightweight Dispersion-Weighted Masking (DWM) strategy that leverages the spectral sparsity of audio spectrograms to dynamically adjust the masking distribution, thereby enhancing representation quality. By prioritizing informative yet sparse regions in the time–frequency domain, DWM significantly reduces computational complexity while consistently improving performance across multiple audio event understanding benchmarks. The approach effectively mitigates the longstanding tension between model efficiency and representational power in audio self-supervised learning.

audio self-supervised learningcomputational overheadgeneralization trade-off

Hot Scholars

ZW

Zongwei Wang

Chongqing University
recommendationdata augmentationpoisoning attack
ZK

Zhifeng Kong

Senior Research Scientist, NVIDIA
Deep Generative ModelsDiffusion ModelsAudio Foundation ModelsAudio LM
XX

Xuenan Xu

Shanghai Jiao Tong University
audio generationaudio understandingspeech synthesis