Beat and Downbeat Tracking in Performance MIDI Using an End-to-End Transformer Architecture

📅 2025-07-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenging tasks of beat and downbeat tracking in symbolic music—specifically, performance-level MIDI—where existing audio-driven approaches dominate. We propose the first end-to-end Transformer sequence-to-sequence model operating entirely in the symbolic domain. Unlike prior methods, our architecture fully leverages an encoder-decoder Transformer, tailored to MIDI’s event-based structure via a dedicated preprocessing pipeline. We further introduce dynamic data augmentation and an optimized symbolic representation scheme to enhance temporal modeling. Evaluated on diverse, multi-instrument datasets—including A-MAPS and ASAP—our method significantly outperforms conventional HMMs and state-of-the-art deep learning baselines, achieving new SOTA F1 scores. Results demonstrate strong generalization across musical styles and robust cross-dataset performance. This work establishes a novel paradigm for symbolic music rhythm analysis, shifting focus from audio-centric to purely symbolic, model-driven approaches.

Technology Category

Cognitive Modeling & Cognitive Systems: Symbolic RepresentationsMachine Learning: Neuro-Symbolic LearningComputer Vision: Visual Reasoning & Symbolic Representations

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
Beat tracking in musical performance MIDI is a challenging and important task for notation-level music transcription and rhythmical analysis, yet existing methods primarily focus on audio-based approaches. This paper proposes an end-to-end transformer-based model for beat and downbeat tracking in performance MIDI, leveraging an encoder-decoder architecture for sequence-to-sequence translation of MIDI input to beat annotations. Our approach introduces novel data preprocessing techniques, including dynamic augmentation and optimized tokenization strategies, to improve accuracy and generalizability across different datasets. We conduct extensive experiments using the A-MAPS, ASAP, GuitarSet, and Leduc datasets, comparing our model against state-of-the-art hidden Markov models (HMMs) and deep learning-based beat tracking methods. The results demonstrate that our model outperforms existing symbolic music beat tracking approaches, achieving competitive F1-scores across various musical styles and instruments. Our findings highlight the potential of transformer architectures for symbolic beat tracking and suggest future integration with automatic music transcription systems for enhanced music analysis and score generation.
Problem

Research questions and friction points this paper is trying to address.

Beat and downbeat tracking in performance MIDI using transformers
Improving accuracy with novel preprocessing and tokenization techniques
Outperforming HMMs and deep learning methods in symbolic beat tracking
Innovation

Methods, ideas, or system contributions that make the work stand out.

End-to-end transformer model for MIDI beat tracking
Dynamic augmentation and optimized tokenization strategies
Outperforms HMMs and deep learning methods
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.