🤖 AI Summary
Existing neural network models struggle to generate cohesive, memorable game music with consistent repetition due to their limited understanding of musical structure, hindering their practical application in gaming contexts. To address this limitation, this work introduces a novel dataset comprising 309 structurally annotated game music audio tracks and proposes the first supervised learning approach for music structure segmentation in this domain. By employing a hybrid CNN-RNN architecture, the model achieves segmentation performance on par with current state-of-the-art unsupervised methods despite using significantly less training data. These results demonstrate the effectiveness and potential of supervised learning for modeling the structural characteristics of game music, offering a promising direction for future research in controllable and structure-aware music generation.
📝 Abstract
At present, neural network-based models, including transformers, struggle to generate memorable and readily comprehensible music from unified and repetitive musical material due to a lack of understanding of musical structure. Consequently, these models are rarely employed by the games industry. It is hypothesised by many scholars that the modelling of musical structure may inform models at a higher level, thereby enhancing the quality of music generation. The aim of this study is to explore the performance of supervised learning methods in the task of structural segmentation, which is the initial step in music structure modelling. An audio game music dataset with 309 structural annotations was created to train the proposed method, which combines convolutional neural networks and recurrent neural networks, achieving performance comparable to the state-of-the-art unsupervised learning methods with fewer training resources.