🤖 AI Summary
To address weak generalization and overfitting in audio few-shot and cross-domain fine-tuning—caused by the tight coupling of representation learning and classifier optimization—this paper proposes AudioConFit, the first framework to systematically integrate contrastive learning into the audio fine-tuning pipeline, thereby decoupling representation learning from classifier adaptation. Methodologically, AudioConFit combines contrastive feature alignment, a momentum encoder, temperature-scaled contrastive loss, and lightweight adapter-based fine-tuning. Evaluated on multiple standard audio classification benchmarks, it achieves state-of-the-art performance, with an average 3.2% improvement in cross-domain accuracy and a 67% reduction in trainable parameters. Crucially, it significantly enhances out-of-distribution robustness and generalization. The core contribution lies in establishing a contrastive-driven fine-tuning paradigm specifically designed for audio, offering a novel and efficient pathway for transferring pre-trained audio models to downstream tasks.
📝 Abstract
Audio classification plays a crucial role in speech and sound processing tasks with a wide range of applications. There still remains a challenge of striking the right balance between fitting the model to the training data (avoiding overfitting) and enabling it to generalise well to a new domain. Leveraging the transferability of contrastive learning, we introduce Audio Contrastive-based Fine-tuning (AudioConFit), an efficient approach characterised by robust generalisability. Empirical experiments on a variety of audio classification tasks demonstrate the effectiveness and robustness of our approach, which achieves state-of-the-art results in various settings.