🤖 AI Summary
This paper addresses automatic music genre classification for multilingual lyrics, proposing a translation-free cross-lingual multi-label genre classification framework. Methodologically, it leverages multilingual Sentence-BERT to encode culturally aware genre semantics from lyrics, integrates One-vs-All classifiers, and employs cross-lingual alignment preprocessing—enabling, for the first time, zero-translation bidirectional (English ↔ Portuguese) genre transfer. A key finding is that dataset centering significantly enhances cross-lingual generalization. Evaluated on a bilingual eight-genre dataset, the framework achieves an average F1-score of 0.69—outperforming a translation-based bag-of-words baseline by 97%—and supports accurate multi-genre annotation per lyric. The framework demonstrates scalability to low-resource languages and adaptability to distinct cultural domains, establishing a transferable semantic modeling paradigm for multilingual music understanding.
📝 Abstract
Music genres are shaped by both the stylistic features of songs and the cultural preferences of artists' audiences. Automatic classification of music genres using lyrics can be useful in several applications such as recommendation systems, playlist creation, and library organization. We present a multi-label, cross-lingual genre classification system based on multilingual sentence embeddings generated by sBERT. Using a bilingual Portuguese-English dataset with eight overlapping genres, we demonstrate the system's ability to train on lyrics in one language and predict genres in another. Our approach outperforms the baseline approach of translating lyrics and using a bag-of-words representation, improving the genrewise average F1-Score from 0.35 to 0.69. The classifier uses a one-vs-all architecture, enabling it to assign multiple genre labels to a single lyric. Experimental results reveal that dataset centralization notably improves cross-lingual performance. This approach offers a scalable solution for genre classification across underrepresented languages and cultural domains, advancing the capabilities of music information retrieval systems.