🤖 AI Summary
Addressing the dual challenges of scarce ground-truth annotations and copyright restrictions on real-world fingerstyle guitar recordings, this paper proposes a knowledge-driven, end-to-end procedural synthesis framework: fingerstyle scores are generated from music-theoretic rules, rendered into MIDI, and converted to high-fidelity audio via an extended Karplus-Strong physical modeling synthesizer, augmented with realistic effects (e.g., reverb, distortion). The resulting fully synthetic dataset enables standalone training of a CRNN-based note-tracking model. When fine-tuned with only a small number of real recordings, the model achieves a significantly higher F1-score on real test data than fully supervised baselines. This work presents the first systematic validation of procedurally generated data for fingerstyle guitar transcription—demonstrating both effectiveness and generalization capability—thereby establishing a reproducible, copyright-compliant, and high-quality data paradigm for low-resource music information retrieval.
📝 Abstract
Automatic transcription of acoustic guitar fingerpicking performances remains a challenging task due to the scarcity of labeled training data and legal constraints connected with musical recordings. This work investigates a procedural data generation pipeline as an alternative to real audio recordings for training transcription models. Our approach synthesizes training data through four stages: knowledge-based fingerpicking tablature composition, MIDI performance rendering, physical modeling using an extended Karplus-Strong algorithm, and audio augmentation including reverb and distortion. We train and evaluate a CRNN-based note-tracking model on both real and synthetic datasets, demonstrating that procedural data can be used to achieve reasonable note-tracking results. Finetuning with a small amount of real data further enhances transcription accuracy, improving over models trained exclusively on real recordings. These results highlight the potential of procedurally generated audio for data-scarce music information retrieval tasks.