π€ AI Summary
This study investigates whether large-scale generative music systems exacerbate musical homogenization and undermine cultural justice. Bridging computational audio analysis with theories of justice for the first time, we conduct a systematic audit of 800 tracks generated by Suno and Lyria using 72-dimensional MIR features alongside metrics of dispersion, redundancy, and separability. Our findings reveal that both systems struggle to faithfully adhere to user prompts, and their outputs are nearly perfectly distinguishable from human-composed music by standard classifiers. Notably, Lyria significantly compresses acoustic diversity within genres, whereas Suno blurs boundaries between genres. These results uncover structurally distinct homogenization patterns across AI systems, posing profound challenges to musical cultural diversity and economic equity.
π Abstract
This paper audits whether large-scale generative music systems exhibit measurable musical homogenization relative to human-produced music, and develops a justice-centered account of why this matters. We audit two commercially deployed systems (Suno and Lyria 3) across four genres (Afrobeats, K-pop, Dance Pop, and Heavy Metal). For each system and genre, we generate 100 tracks and compare them against human corpora of equal size, using 72 music information retrieval (MIR) features and multiple diagnostics of dispersion, redundancy, and separability. We define homogenization as reduced acoustic variation in standard computational audio features including rhythm and timing, timbre/spectral shape, and dynamics, both within genres and across genre boundaries. We also generate tracks using only a genre name as the prompt, with no additional instructions, to reveal each system's default musical tendencies. The results show two structurally distinct homogenizing tendencies. Lyria reduces within-genre acoustic diversity, while Suno collapses the acoustic distinctions between genres without compressing within-genre spread. Neither system follows user prompts faithfully, indicating that the observed patterns reflect learned priors rather than prompt constraints. The two systems do not converge on a common acoustic profile and are more acoustically distant from each other than two random human subsamples would typically be. Nevertheless, a standard classifier distinguishes AI from human tracks near-perfectly on MIR features alone. We argue that these patterns matter not as an aesthetic curiosity but as a justice-relevant condition, shaping which musical styles become legible, valued, and economically rewarded as generated outputs increasingly circulate at scale.