A Survey on Evaluation Metrics for Music Generation

📅 2025-08-24
📈 Citations: 0
Influential: 0
📄 PDF

career value

199K/year
🤖 AI Summary
Evaluation of music generation lags behind model development due to weak correlation between objective metrics and human perception, significant cross-cultural biases, and the absence of standardized benchmarks. Method: Through a systematic literature review and critical analysis, we integrate methodologies from music information retrieval, audio signal processing, and user experience evaluation to construct the first taxonomy of multidimensional evaluation metrics covering both audio and symbolic music representations. Contribution/Results: Our analysis uncovers fundamental deficiencies in existing approaches—namely, low inter-subjective consistency, poor cultural universality, and limited reproducibility—and identifies three core bottlenecks: insufficient perceptual alignment, lack of cultural sensitivity, and fragmented benchmarking practices. The proposed taxonomy establishes a theoretical foundation and practical roadmap for developing standardized, comparable, and human-centered evaluation frameworks, thereby advancing the scientific rigor and methodological coherence of music generation assessment.

Technology Category

Application Category

📝 Abstract
Despite significant advancements in music generation systems, the methodologies for evaluating generated music have not progressed as expected due to the complex nature of music, with aspects such as structure, coherence, creativity, and emotional expressiveness. In this paper, we shed light on this research gap, introducing a detailed taxonomy for evaluation metrics for both audio and symbolic music representations. We include a critical review identifying major limitations in current evaluation methodologies which includes poor correlation between objective metrics and human perception, cross-cultural bias, and lack of standardization that hinders cross-model comparisons. Addressing these gaps, we further propose future research directions towards building a comprehensive evaluation framework for music generation evaluation.
Problem

Research questions and friction points this paper is trying to address.

Evaluating generated music lacks standardized metrics
Poor correlation between objective metrics and human perception
Cross-cultural bias and lack of standardization hinder comparisons
Innovation

Methods, ideas, or system contributions that make the work stand out.

Proposed taxonomy for audio and symbolic metrics
Identified limitations in current evaluation methodologies
Suggested future directions for comprehensive framework