🤖 AI Summary
This work addresses the limitations of existing collaborative music-generation agents, which lack internal representations that jointly support understanding and generation while preserving human creative control. We propose a hierarchical self-supervised world model trained on unlabeled MIDI piano rolls using a Swin V2 encoder and a JEPA-style objective to learn musical structure. High-quality generation and interactive prompting are achieved via conditional flow matching. Notably, our approach spontaneously discovers temporal and phrase-level structures without prior music-theoretic knowledge, enhances interpretability through hierarchical embeddings, and enables graphical mask-based inpainting without specialized samplers. Experiments show that frozen embeddings disentangle musical attributes across timescales; adding a lightweight chord-supervision head boosts chord recovery accuracy from 0.18 to 0.54 and achieves a tonality detection score of 0.70. The model attains a generation F1 score of 0.996, with inference times of 2.8 seconds on CPU and just 0.6 seconds on Apple MPS, and has been integrated into a real-time interactive system.
📝 Abstract
Collaborative music agents need internal representations rich enough to support both understanding and generation, yet flexible enough for a workflow where the human retains agency. We present a hierarchical self-supervised ``world model'' for symbolic music: a 2.55M-parameter Swin V2 encoder trained on MIDI piano-roll images with JEPA-style objectives (pitch- and time-shift equivariance, masked embedding prediction, and a distributional regularizer), using no labels and no music-theory vocabulary. Probing the frozen embeddings shows that the level at which a musical property becomes decodable tracks its musical time scale: phrase boundaries are read off the coarsest levels, note density and harmonic detail off the finest. Temporal and phrase structure emerge from the self-supervised objectives alone, while harmonic content must be asked for; a small chord-supervision head raises joint chord recovery from .18 to .54, and key detection, which is never supervised, from .16 to .70. Following the Representation AutoEncoder paradigm, a conditional flow-matching model stands in for a trained decoder, flowing in pixel space from PCA-reduced conditioning: it reproduces a target window at pixel F1 $0.996$, and the same per-level conditioning dropout that controls how far variations stray also enables graphical prompting for masked inpainting with no inpainting-specific sampler. The pipeline runs on CPU producing a suggestion in $2.8$ s, or $0.6$ s on Apple MPS, which we demonstrate in a live interactive demo. In concert with an LLM-based brain, these capabilities supply the core of a collaborative music creation agent in service of, rather than in place of, human agency.