🤖 AI Summary
This study addresses the challenge that existing systems struggle to disentangle timbre from melody and genre in audio references, thereby limiting zero-shot multi-instrument timbre transfer. To overcome the constraints of text prompting, this work proposes the first zero-shot polyphonic timbre transfer framework that conditions a pretrained music generator on audio references. Specifically, a fine-tuned CLAP encoder extracts reference timbre features while a multi-pitch estimator parses source melodies, enabling precise disentanglement of timbre and content within a frozen generative model. Evaluated across four real-world polyphonic datasets, the proposed method significantly improves timbre matching fidelity while preserving pitch alignment accuracy. Furthermore, the feature extractor demonstrates strong robustness to pitch variations, establishing an effective approach for high-quality, controllable multi-instrument timbre transfer without requiring paired training data or textual descriptions.
📝 Abstract
Instrument timbre transfer re-voices a performance using the timbre of another instrument. Extracting the target timbre from an audio reference capture more nuances than inferring it from a text prompt. Systems that read timbre from such a clip train a dedicated model for the task, which captures the timbre cleanly but stays a narrow, single-purpose system. More versatile approaches add control to a pretrained music generator, yet a reference clip entangles timbre with genre and melody, so these systems fall back on text to name the timbre. We present MuseTimbre, the first system, to our best knowledge, that transfers timbre from an audio reference to a polyphonic source through conditioning a pretrained music generator. This system employs a multi-pitch estimator to extract pitch information from the source and finetune a CLAP encoder to extract timbre information from the reference audio. Experiments show that across four datasets of real polyphonic recordings, MuseTimbre achieves pitch alignment on par with the baselines while matching the reference timbre far more closely. Results also show that the finetuned CLAP-based timbre extractor is robust to pitch variations, making it useful in timbre similarity measures.