Extending Music Annotation Schemas: Zero-Shot Prediction or Few-Shot Adaptation?

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the trade-off between annotation costs for zero-shot prediction and supervised adaptation when extending new attributes to music catalogs. To simulate schema expansion scenarios, we construct the MGPHot benchmark and systematically compare audio-language models across zero-shot inference, pretrained representation reuse, and few-shot supervised fine-tuning strategies. This work is the first to quantify performance disparities among these methods under varying annotation budgets. Our findings demonstrate that supervised adaptation significantly outperforms zero-shot prediction in low-budget settings, while frozen representation reuse emerges as the optimal solution under moderate budgets without requiring extensive tuning. Ultimately, this research provides practical, evidence-based guidance for strategy selection during dynamic attribute expansion in music information retrieval systems.
📝 Abstract
Automatic music annotation is typically tackled under the assumption of a fixed annotation schema. In practice, commercial music catalogs often need to accommodate new musical attributes as needs evolve. Given that expert music annotation is expensive, it is not evident which methodological approach is most effective at accommodating new attributes and backfilling existing tracks; audio-language models promise zero-shot prediction, but at what annotation budget does supervised adaptation become more compelling? We propose a benchmark based on the MGPHot popular music annotation dataset for simulating music schema extension across different annotation budgets. We investigate zero-shot prediction with audio-language models, learning new attributes from pretrained representations, and adapting models trained on existing annotations. Our results suggest that supervised adaptation is more effective than zero-shot prediction even with small annotation budgets, while frozen representation reuse remains the most effective approach for modest budgets without the tuning required by deeper adaptation.
Problem

Research questions and friction points this paper is trying to address.

music annotation
schema extension
zero-shot prediction
few-shot adaptation
audio-language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Zero-Shot Prediction
Few-Shot Adaptation
Music Annotation Schema
Audio-Language Models
Pretrained Representations
🔎 Similar Papers
2024-09-17IEEE Open Journal of Signal ProcessingCitations: 0