🤖 AI Summary
This study systematically compares three major classes of generative models—autoregressive LSTMs with attention, latent-variable models (including recurrent VAEs and VQ-VAEs), and GANs—in the task of generating polyphonic symbolic music in the style of Bach. Evaluated on a unified MIDI benchmark within a common framework, the models are assessed for musical coherence, quality of learned latent representations, and stylistic consistency. Results indicate that autoregressive LSTMs produce the most coherent sequences; VQ-VAEs effectively mitigate posterior collapse through vector quantization, yielding clearer structural outputs; and GANs, while capable of capturing local pitch patterns, suffer from training instability and poor generalization. The findings elucidate the respective strengths and failure modes of each approach, offering empirical guidance for polyphonic music generation.
📝 Abstract
We study generative modeling of Bach-style symbolic piano music using a shared MIDI corpus and three model families: autoregressive LSTMs with attention, latent-variable models including recurrent VAEs and vector-quantized VAEs, and generative adversarial networks. We compare their ability to model polyphonic note sequences, learn useful latent representations, and generate stylistically coherent compositions. Our experiments show that the autoregressive LSTM with attention produces the most musically coherent samples, while vector quantization helps mitigate posterior collapse and yields more structured outputs than conventional recurrent VAEs. The adversarial approach captures local pitch patterns but remains difficult to train and generalizes less reliably to Bach's style. These results highlight the relative strengths and failure modes of autoregressive, latent-variable, and adversarial approaches for symbolic music generation.