Learning task-specific subspaces via interventional post-training of speech foundation models

📅 2026-06-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the entanglement of linguistic content and speaker information in general-purpose representations from speech foundation models, which limits performance on downstream tasks requiring only one of these factors. To resolve this without re-pretraining, the authors propose an interventional contrastive learning approach for post-training optimization. By constructing an interventional dataset and designing a multipart contrastive loss, the method explicitly disentangles the representation into independent subspaces for content and speaker identity. This study presents the first application of interventional contrastive learning to speech representation disentanglement, achieving significant performance gains on cross-domain speaker verification tasks. Experimental results demonstrate that the learned subspaces effectively separate the two semantic attributes, confirming the efficacy of the proposed disentanglement strategy.
📝 Abstract
Speech foundation models, pre-trained on large corpora of unlabelled speech data, produce general-purpose representations which are useful across tasks. However, these representations encode information about salient speech variables in a distributed manner, while downstream speech tasks rely on only some of this variability. In this work, we propose a post-training refinement approach using interventional contrastive learning. By leveraging an interventional dataset and multi-part contrastive loss, we learn a transformation from the entangled representation space of speech foundation models into separate content and speaker subspaces. We evaluate the learnt representations on speaker verification and keyword spotting tasks, showing improved out-of-domain speaker verification performance and evidence that speaker and content information are separated across the learned subspaces.
Problem

Research questions and friction points this paper is trying to address.

speech foundation models
representation disentanglement
task-specific subspaces
speaker-content separation
downstream speech tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

interventional contrastive learning
speech foundation models
representation disentanglement
post-training refinement
subspace learning
💼 Related Jobs
No related jobs found.
J
Jack Cox
University of Sheffield, United Kingdom
J
Jon Barker
University of Sheffield, United Kingdom