π€ AI Summary
This study investigates semantic alignment between user prompts and listener perceptions in text-to-music generation, with a particular focus on how linguistic differences across cultures influence this alignment. Leveraging 200 promptβaudio pairs generated by Udio, the authors collected free-text descriptions from English- and Korean-speaking participants to construct, for the first time, a human-annotated taxonomy of musical prompting language. Employing both word-level and vector-level semantic analyses, they evaluate cross-lingual alignment and find that prompts are predominantly genre- and narrative-driven, with genre-related terms most accurately perceived, whereas highly narrative content often leads to semantic misalignment. Furthermore, significant cross-cultural differences emerge in how listeners describe music along narrative, functional, and emotional dimensions.
π Abstract
Text-to-music (TTM) generation systems allow users to create music through natural language prompts, yet it is unclear whether the descriptive language used to prompt aligns with descriptive language used to summarize or describe heard music. We pair 200 real-world Udio prompts with their generated audio and free-form descriptions collected from English- (n = 70) and Korean-speaking (n = 78) listeners, and contribute a human-derived taxonomy of musical prompting vocabulary grounded in real user data. Using this framework, alongside word- and vector-level analyses, we find a consistent structural asymmetry: prompts are dominated by Genre and Story/Narrative language. Genre terms propagate most reliably from prompt to perception, while narrative-heavy prompts are the strongest predictor of semantic misalignment. A preliminary cross-cultural comparison further suggests that description profiles vary across listener populations along narrative, functional, and affective dimensions, raising questions about whether current TTM systems, trained on aggregated English-centric corpora, can accommodate the full diversity of how people naturally express musical ideas.