Synthesis and editing of multi-instrument audio mixtures using scalar-quantised latents with MIDI Span conditioning

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
该研究提出SpanSynth-Edit模型,利用MIDI Span条件和低帧率标量量化潜变量,支持多乐器音频混合的合成与编辑。
📝 Abstract
Music creation often involves iterative refinement, changing selected musical details while retaining the rest. To support such refinement, we introduce SpanSynth-Edit, a flow-matching model for MIDI-guided synthesis and editing of multi-instrument audio mixtures using low-frame-rate scalar-quantised latents. MIDI Span encodes instrument-labelled note lifecycles as unordered event sets with continuous-valued attributes and pools each set into one conditioning vector per audio-latent frame. The model uses contextual audio for instrument-specific timbre guidance and supports editing by resynthesising the target region from revised MIDI. Experiments on single- and multi-instrument benchmarks show competitive performance and demonstrate within-frame onset control. We also discuss limitations of transcription-based note-adherence evaluation.
Problem

Research questions and friction points this paper is trying to address.

MIDI-guided synthesis
multi-instrument audio mixtures
iterative refinement
Innovation

Methods, ideas, or system contributions that make the work stand out.

flow-matching model
MIDI-guided synthesis
scalar-quantised latents
multi-instrument audio mixtures
contextual audio
🔎 Similar Papers