🤖 AI Summary
Existing text-controlled symbolic music generation methods suffer from limitations in generation quality, diversity, controllability, and output duration. This work proposes a novel framework that, for the first time, integrates autoregressive modeling with a latent diffusion model (LDM) for this task. The approach introduces an efficient music information encoder to enhance textual control fidelity and leverages large language models to construct the first large-scale paired text–symbolic music dataset. Experimental results demonstrate that the proposed method significantly outperforms strong baselines—including GPT-4, MuseCoco, and MMT—in terms of audio quality, generation length, and controllability, establishing a new state of the art in text-to-music synthesis.
📝 Abstract
Text-controlled symbolic music generation has recently gained research attention due to its versatile, flexible and straightforward approach to music composition. However, previous approaches tend to generate symbolic music with compromising quality, diversity, controllability and limited duration. In this paper, we present Diff-Symbo, an innovative method that uses latent diffusion model (LDM) to generate high-quality, diverse and long-duration symbolic music. To address the lack of text-symbolic music dataset, we develop a comprehensive dataset with 19,345 text templates by employing large language model. Furthermore, we design a music information encoder to reduce the training overhead while extracting more effective control representations. Given textual descriptions, our proposed method leverages LDM to improve the quality and diversity of music generation. Our method also improves the duration and the compositional consistency of music generation through an autoregressive approach. Experimental results show significant improvements of Diff-Symbo in text controllability, duration, and the quality of generated music compared to the baseline models such as GPT-4, MuseCoco and Multitrack Music Transformer (MMT). As one of the pioneer models in this field, Diff-Symbo paves the way towards controllable and high-quality symbolic music composition based on LDM, offering valuable contributions to both music amateurs and practitioners.