🤖 AI Summary
This work addresses the challenging tasks of beat and downbeat tracking in symbolic music—specifically, performance-level MIDI—where existing audio-driven approaches dominate. We propose the first end-to-end Transformer sequence-to-sequence model operating entirely in the symbolic domain. Unlike prior methods, our architecture fully leverages an encoder-decoder Transformer, tailored to MIDI’s event-based structure via a dedicated preprocessing pipeline. We further introduce dynamic data augmentation and an optimized symbolic representation scheme to enhance temporal modeling. Evaluated on diverse, multi-instrument datasets—including A-MAPS and ASAP—our method significantly outperforms conventional HMMs and state-of-the-art deep learning baselines, achieving new SOTA F1 scores. Results demonstrate strong generalization across musical styles and robust cross-dataset performance. This work establishes a novel paradigm for symbolic music rhythm analysis, shifting focus from audio-centric to purely symbolic, model-driven approaches.
📝 Abstract
Beat tracking in musical performance MIDI is a challenging and important task for notation-level music transcription and rhythmical analysis, yet existing methods primarily focus on audio-based approaches. This paper proposes an end-to-end transformer-based model for beat and downbeat tracking in performance MIDI, leveraging an encoder-decoder architecture for sequence-to-sequence translation of MIDI input to beat annotations. Our approach introduces novel data preprocessing techniques, including dynamic augmentation and optimized tokenization strategies, to improve accuracy and generalizability across different datasets. We conduct extensive experiments using the A-MAPS, ASAP, GuitarSet, and Leduc datasets, comparing our model against state-of-the-art hidden Markov models (HMMs) and deep learning-based beat tracking methods. The results demonstrate that our model outperforms existing symbolic music beat tracking approaches, achieving competitive F1-scores across various musical styles and instruments. Our findings highlight the potential of transformer architectures for symbolic beat tracking and suggest future integration with automatic music transcription systems for enhanced music analysis and score generation.