🤖 AI Summary
This study addresses the risk of artist identity leakage in lyrics-to-song generative models, where lyrics serve as an unguarded conditioning channel that may expose training data privacy. To investigate this, we conduct a systematic model audit on 2,000 songs by probing the internal activations of ACE-Step 1.5 using linear probes and latent space analysis techniques. We provide the first demonstration that artist identity representations can be linearly decoded from internal model activations based solely on input lyrics, and reveal the propagation mechanism of this signal from the lyrics encoder to the diffusion backbone. By successfully identifying potential artist associations, this work bridges a critical gap in interpretability research within audio generation and validates latent space analysis as an effective approach for auditing implicitly learned content in generative models.
📝 Abstract
Text-to-song generation models can be prompted to imitate specific artists or regurgitate entire songs from their training data. Although these phenomena have been documented behaviorally on small datasets, little is known about the internal representations that may give rise to them. Prior interpretability work on generative audio has focused on locating semantic concepts such as genre or time signature within model activations. In this work, we show that a trained model can be probed for linearly decodable representations of artist identity from song lyrics alone, without any additional identifiers. Through a controlled case study of ACE-Step 1.5 spanning 2,000 songs across 100 artists, we demonstrate that the artist associated with a given set of lyrics can be identified within the model's internal activations, and that this conditioning signal propagates from the lyric encoder to the diffusion backbone during inference. These findings indicate that lyrics constitute an artist-level conditioning channel not addressed by prompt-side replication safeguards. More broadly, our work highlights how latent-space analysis can be used to audit what generative music models have implicitly learned from their training data.