🤖 AI Summary
This paper addresses the problem of automatic MIDI performance quantization into canonical musical notation. Methodologically, it proposes a Transformer model that jointly incorporates beat structure and accent perception: (i) a beat-aligned preprocessing pipeline unifies symbolic representations of performances and scores; (ii) beat-position encoding and accent-aware token embeddings explicitly model hierarchical metrical relationships. Trained on multi-style piano and guitar performance data, the model achieves state-of-the-art performance on the MUSTER benchmark—improving rhythmic quantization accuracy by 4.2 percentage points (absolute) over prior methods, with notably enhanced robustness on complex rhythms and freely interpreted passages. The core contribution lies in deeply integrating beat-level structural priors into the Transformer architecture, enabling end-to-end, interpretable, high-accuracy music quantization.
📝 Abstract
We propose a transformer-based rhythm quantization model that incorporates beat and downbeat information to quantize MIDI performances into metrically-aligned, human-readable scores. We propose a beat-based preprocessing method that transfers score and performance data into a unified token representation. We optimize our model architecture and data representation and train on piano and guitar performances. Our model exceeds state-of-the-art performance based on the MUSTER metric.