Fine-Tuning MIDI-to-Audio Alignment using a Neural Network on Piano Roll and CQT Representations

📅 2025-06-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the low synchronization accuracy between piano performance audio and loosely aligned MIDI files. We propose an end-to-end temporal alignment framework based on a Convolutional Recurrent Neural Network (CRNN). The method jointly processes piano roll representations and Constant-Q Transform (CQT) spectrograms to jointly model time-frequency and long-range temporal structural features; it incorporates a performance error simulation strategy for data augmentation and integrates a Dynamic Time Warping (DTW)-based post-processing module to enhance robustness. On standard benchmark datasets, our approach achieves up to a 20% improvement in alignment accuracy across multiple tolerance thresholds (10–500 ms), significantly outperforming conventional DTW and state-of-the-art deep learning methods. Key contributions include: (i) the first systematic application of CRNNs to fine-grained MIDI-audio alignment; (ii) empirical validation that neural networks effectively capture performance-induced temporal variability; and (iii) simultaneous enhancement of both alignment precision and robustness.

Technology Category

Planning, Routing, and Scheduling: Temporal PlanningMachine Learning: Matrix & Tensor MethodsCognitive Modeling & Cognitive Systems: Neural Spike Coding

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
In this paper, we present a neural network approach for synchronizing audio recordings of human piano performances with their corresponding loosely aligned MIDI files. The task is addressed using a Convolutional Recurrent Neural Network (CRNN) architecture, which effectively captures spectral and temporal features by processing an unaligned piano roll and a spectrogram as inputs to estimate the aligned piano roll. To train the network, we create a dataset of piano pieces with augmented MIDI files that simulate common human timing errors. The proposed model achieves up to 20% higher alignment accuracy than the industry-standard Dynamic Time Warping (DTW) method across various tolerance windows. Furthermore, integrating DTW with the CRNN yields additional improvements, offering enhanced robustness and consistency. These findings demonstrate the potential of neural networks in advancing state-of-the-art MIDI-to-audio alignment.
Problem

Research questions and friction points this paper is trying to address.

Synchronizing piano audio with loosely aligned MIDI files
Improving alignment accuracy using neural networks
Combining CRNN and DTW for robust MIDI-audio alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

CRNN processes piano roll and spectrogram inputs
Augmented MIDI dataset simulates human timing errors
Combines CRNN and DTW for improved alignment accuracy
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Sebastian Murgul
Sebastian Murgul
Institute of Industrial Information Technology, Karlsruhe Institute of Technology, Karlsruhe
Music Information RetrievalAutomatic Music TranscriptionMachine LearningArtificial Intelligence
M
Moritz Reiser
University of Music Karlsruhe
M
Michael Heizmann
Karlsruhe Institute of Technology
C
Christoph Seibert
University of Music Karlsruhe