Towards Robust Version Identification in the Wild: A Dataset, Benchmark, and Fine-Tuning Study

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the significant domain mismatch in existing music version identification datasets, which predominantly rely on professional recordings and fail to generalize to real-world amateur or user-generated content. To bridge this gap, the authors introduce DiVers, a large-scale dataset comprising over 1.1 million annotated music version samples that explicitly incorporate real-world audio diversity and noise. DiVers is compatible with multiple established benchmarks and leverages automated annotation techniques to provide rich semantic labels—including instrumentation and performance environment—as well as precise music segment boundaries, enabling fine-grained analysis. Experimental results demonstrate that models trained on DiVers achieve substantially improved robustness on noisy and diverse audio while maintaining consistent performance on clean, professionally recorded benchmarks.
📝 Abstract
Existing datasets for musical version identification (VI) are primarily derived from curated metadata sources such as SecondHandSongs and Discogs, and are therefore dominated by professionally recorded tracks. This leads to a domain mismatch with real-world scenarios, where amateur and user-generated content is prevalent. To address this limitation, we introduce DiVers, a large-scale VI dataset comprising over 1.1 million musical versions, with train-validation-test splits compatible with established datasets such as Discogs-VI-YT, SHS100K, and Da-TACOS. In addition to standard version-level annotations, DiVers provides automatically assigned tags (e.g., instrumental, live) and segment-level predictions indicating the presence or absence of music. We evaluate the proposed dataset by training state-of-the-art VI systems. Our results show that models trained on DiVers achieve substantially improved robustness to acoustically diverse and noisy inputs, while maintaining a stable performance on cleaner, studio-quality benchmarks. We release the dataset metadata, code for its construction, and all experimental pipelines to support reproducibility.
Problem

Research questions and friction points this paper is trying to address.

version identification
domain mismatch
user-generated content
robustness
music dataset
Innovation

Methods, ideas, or system contributions that make the work stand out.

version identification
robustness
user-generated content
large-scale dataset
domain generalization