🤖 AI Summary
This work addresses the lack of guitar-specific physical constraints—such as string-fret mapping and playability—in MIDI symbolic music. We propose the first T5-based method for automatic transcription from MIDI to guitar tablature, framing the task as conditional sequence-to-sequence translation. Our approach innovatively jointly models string-fret ambiguity resolution and physical playability constraints, while incorporating tunings and capo configurations as controllable conditions and employing context-sensitive decoding. We introduce a customized tokenization scheme, train on fused multi-source datasets (DadaGP, GuitarToday, Leduc), and design a dual-axis evaluation metric combining fingering accuracy and playability. Experiments demonstrate significant improvements over A*-search baselines and commercial tools (e.g., Guitar Pro) across all metrics. The method supports arbitrary tunings and capo positions, generating tablatures that are both highly accurate and practically playable.
📝 Abstract
Music transcription plays a pivotal role in Music Information Retrieval (MIR), particularly for stringed instruments like the guitar, where symbolic music notations such as MIDI lack crucial playability information. This contribution introduces the Fretting-Transformer, an encoderdecoder model that utilizes a T5 transformer architecture to automate the transcription of MIDI sequences into guitar tablature. By framing the task as a symbolic translation problem, the model addresses key challenges, including string-fret ambiguity and physical playability. The proposed system leverages diverse datasets, including DadaGP, GuitarToday, and Leduc, with novel data pre-processing and tokenization strategies. We have developed metrics for tablature accuracy and playability to quantitatively evaluate the performance. The experimental results demonstrate that the Fretting-Transformer surpasses baseline methods like A* and commercial applications like Guitar Pro. The integration of context-sensitive processing and tuning/capo conditioning further enhances the model's performance, laying a robust foundation for future developments in automated guitar transcription.