🤖 AI Summary
This study addresses the absence of reliable benchmarks and ground truth for data attribution in text-to-music generation models by constructing the first rigorous evaluation benchmark. By fine-tuning diffusion models to synthesize controlled datasets, this work provides known attribution targets across multiple dimensions, including melody and timbre. These targets are integrated with black-box attribution algorithms to establish a large-scale, multi-task evaluation framework. Experimental results reveal that existing attribution methods exhibit substantial performance discrepancies across diverse scenarios and remain highly challenging, underscoring the urgent need to develop generalizable and robust attribution techniques. Ultimately, this work provides the field with a standardized testing platform for systematically assessing and advancing data attribution methodologies in generative audio models.
📝 Abstract
Text-to-music generation models are trained on massive music collections, creating a growing need for data attribution methods that can quantify the contribution of individual training samples. However, existing attribution methods are difficult to rigorously evaluate due to the lack of reliable ground truth, making it challenging to reliably assess their actual effectiveness. To address this gap, we introduce TrueMuse, a controlled dataset and benchmark for text-to-music data attribution. TrueMuse is constructed by fine-tuning three diffusion-based text-to-music models on carefully curated attribution samples, whose known inclusion in fine-tuning provides controlled attribution targets for evaluation. The benchmark covers four attribution settings, spanning melodic structure, timbral characteristics, artist-level stylistic signatures, and genre-level shared patterns, and includes 133 attributes, 648 fine-tuned models, and 95,456 generated samples across two prompt types. Using TrueMuse, we systematically evaluate existing black-box attribution methods along four dimensions: fine-tuning improvement, prompt-type difficulty, multi-task training, and fine-tuning data size. Our results show that attribution remains challenging, with existing methods exhibiting substantial variation across evaluation settings, highlighting the need for more reliable and generalizable attribution methods for text-to-music generation. Code and Dataset will be released upon acceptance.