Learning Jazz Pianist Style with Cross-Attention Conditioning
This study addresses the challenge of formalizing jazz piano performance styles and generating music conditioned on specific artists. The proposed method builds upon a pretrained symbolic music Transformer, injecting learnable pianist identity embeddings via cross-attention mechanisms to achieve style-conditioned generation. Furthermore, it pioneers the integration of pretrained representations with a sliding-window classifier for style feature localization and classification. The framework effectively captures latent stylistic structures, significantly outperforming baseline models in conditional generation. Notably, classifiers trained exclusively on synthetic data attain 87% segment-level and 95% song-level accuracy in real-world pianist identification tasks, demonstrating the model’s capacity to synthesize authentic stylistic characteristics and its potential for data-efficient style analysis.