π€ AI Summary
This study addresses the challenge of formalizing jazz piano performance styles and generating music conditioned on specific artists. The proposed method builds upon a pretrained symbolic music Transformer, injecting learnable pianist identity embeddings via cross-attention mechanisms to achieve style-conditioned generation. Furthermore, it pioneers the integration of pretrained representations with a sliding-window classifier for style feature localization and classification. The framework effectively captures latent stylistic structures, significantly outperforming baseline models in conditional generation. Notably, classifiers trained exclusively on synthetic data attain 87% segment-level and 95% song-level accuracy in real-world pianist identification tasks, demonstrating the modelβs capacity to synthesize authentic stylistic characteristics and its potential for data-efficient style analysis.
π Abstract
Jazz pianists develop distinctive traits that experienced listeners can often identify within seconds, yet the features underlying this recognition resist formal description. We study jazz pianist style through the lens of a pretrained symbolic music transformer, showing that its learned representations already encode pianist identity well enough for highly accurate classification across two benchmarks. We then augment the transformer with cross-attention over learned pianist identity embeddings, enabling it to generate music conditioned on a specific artist's style. Two evaluation protocols confirm that the generator captures meaningful stylistic structure: a sliding-window classifier consistently attributes conditioned continuations to the correct artist, far above unconditioned baselines; and a classifier trained entirely on synthetic generations identifies real pianists across 12 classes with 87% chunk-level and 95% song-level accuracy. Finally, we repurpose the classifier to locate the most characteristic moments within a performance, surfacing the specific musical gestures that distinguish each pianist's voice.