PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing

📅 2024-09-17
🏛️ IEEE International Conference on Acoustics, Speech, and Signal Processing
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Generative AI for music faces core challenges including copyright restrictions on symbolic music data, scarcity of high-quality public-domain resources, and difficulties in objective quality assessment. To address these, we introduce PDMX—the largest open-source, copyright-free MusicXML dataset to date—comprising over 250,000 public-domain scores, systematically augmented with interactive metadata such as user ratings and performance feedback. This work pioneers the use of user behavioral data for symbolic music quality modeling; empirical analysis demonstrates its strong predictive power for generation quality. Furthermore, we reveal that subset selection critically influences the performance of polyphonic generation models. The dataset is publicly released and enables more stable, reproducible training of downstream models across melody, harmony, and rhythm modeling tasks.

Technology Category

Natural Language Processing: GenerationMachine Learning: Large Multimodal Models (LMMs)Humans and AI: Game Design — Procedural Content Generation & Storytelling

Application Category

Economics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsWeb Mining and Content Analysis: Web data quality in the era of algorithmically-generated contentUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
The recent explosion of generative AI-Music systems has raised numerous concerns over data copyright, licensing music from musicians, and the conflict between open-source AI and large prestige companies. Such issues highlight the need for publicly available, copyright-free musical data, in which there is a large shortage, particularly for symbolic music data. To alleviate this issue, we present PDMX: a large-scale open-source dataset of over 250K public domain MusicXML scores collected from the score-sharing forum MuseScore, making it the largest available copyright-free symbolic music dataset to our knowledge. PDMX additionally includes a wealth of both tag and user interaction metadata, allowing us to efficiently analyze the dataset and filter for high quality user-generated scores. Given the additional metadata afforded by our data collection process, we conduct multitrack music generation experiments evaluating how different representative subsets of PDMX lead to different behaviors in downstream models, and how user-rating statistics can be used as an effective measure of data quality. Examples can be found at https://pnlong.github.io/PDMX.demo/.
Problem

Research questions and friction points this paper is trying to address.

Addresses the shortage of copyright-free symbolic music data.
Provides a large-scale open-source MusicXML dataset (PDMX).
Explores multitrack music generation using metadata for quality filtering.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large-scale public domain MusicXML dataset
Includes tag and user interaction metadata
Multitrack music generation experiments conducted
UC San Diego
P
Phillip Long
Computer Science & Engineering, UC San Diego
Zachary Novack
Zachary Novack
CS PhD Student, UC - San Diego
music generationdiffusion modelsmultimodality
Taylor Berg-Kirkpatrick
Taylor Berg-Kirkpatrick
University of California San Diego
Natural Language Processing
J
Julian J. McAuley
Computer Science & Engineering, UC San Diego