🤖 AI Summary
Existing self-supervised methods struggle to effectively model the hierarchical structure of symbolic music, limiting performance in tasks such as machine-assisted composition. This work proposes a hierarchical self-supervised model that, for the first time, incorporates temporal and pitch translation equivariance into symbolic music representation learning. Built upon a Swin Transformer V2 encoder within the LeJEPA framework, the model combines masked embedding prediction with SIGReg regularization to prevent representational collapse. The learned representations exhibit both semantic richness and generative utility: decoded reconstructions achieve an F1 score of 0.995; generated music precisely matches the pitch range and rhythmic density of conditioning fragments; downstream emotion classification outperforms Haar scattering baselines; and embedding distances vary monotonically with pitch or temporal shifts, confirming the efficacy of equivariant modeling.
📝 Abstract
Rich internal representations of musical structure are essential for music understanding tasks such as machine-assisted music co-writing, yet self-supervised approaches for symbolic music representation remain underexplored, particularly those that encode the hierarchical multiscale nature of musical structures. We present MIDI-RAE-JEPA, combining a pitch- and time-shift equivariance objective with LeJEPA and a Swin Transformer V2 encoder to learn such hierarchical representations of symbolic music encoded as piano roll images. The time-shift equivariance objective encourages the model to internalize temporal musical relationships. The encoder is trained purely on self-supervised objectives -- including a masked embedding predictor (MEP) -- with collapse prevented via SIGReg. A separate decoder trained on the frozen encoder embeddings achieves reconstruction F1 of 0.995, and a flow matching generative model conditioned on those embeddings produces generations that closely match the pitch register and rhythmic density of the conditioning excerpt, while mismatched conditioning yields unrelated but musically plausible output. Learned representations outperform a Haar scattering transform baseline on a downstream emotion classification task, and embedding distances increase monotonically with pitch and time shift magnitude, confirming measurable equivariance. These results suggest that equivariance-based SSL objectives, combined with sufficient fine-level encoder capacity, provide a viable path toward semantically rich, generatively useful representations of symbolic music.