Retrieving Individual Stems from Music Mixtures with Slot Embeddings

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing music retrieval systems that rely on instrument labels and struggle to isolate target tracks from mixed audio. To this end, we propose Stembed, a model that introduces a novel slot embedding-based representation mechanism for mixed audio. By leveraging deep neural networks and slot attention, Stembed encodes mixed signals into multiple candidate slots. Contrastive learning pairs are constructed from tracks of the same song, enabling slots to automatically inherit track identity and facilitating precise retrieval without prior information. Experimental results on the MoisesDB dataset demonstrate that the proposed method comprehensively outperforms composed music information retrieval (CIR) baselines. Notably, even without instrument-based filtering, Stembed achieves significant improvements in the R@1 metric, validating its effectiveness for robust music retrieval in complex acoustic mixtures.
📝 Abstract
Music producers search libraries of isolated instrument recordings, called stems, for sounds resembling parts of an existing song. Neural retrieval systems address this by mapping audio to embeddings and ranking library stems by their similarity to the query. The leading method, Contrastive Instrument Retrieval (CIR), encodes the mixture as a single embedding, but it works best when a user specifies the target's instrument family. We introduce Stembed, which encodes a mixture as several slot embeddings representing candidate stems. During training, we construct mixtures from stems of the same song and match their slot embeddings to those of the isolated stems. The slot embeddings from mixtures inherit the stem identities of their assigned solo embedding, enabling a contrastive loss. On mixtures from held out MoisesDB artists, Stembed outperforms a CIR-style baseline when both search the full stem library. Even when predicting the stem count itself without family labels, Stembed exceeds the baseline's family-filtered R@1. Our website demonstrates how users can select a slot by inspecting the tags of its retrieved stems.
Problem

Research questions and friction points this paper is trying to address.

music information retrieval
source separation
slot embeddings
contrastive learning
audio representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Slot Embeddings
Contrastive Learning
Music Information Retrieval
Source Separation
Audio Representation
🔎 Similar Papers
No similar papers found.