YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

πŸ“… 2026-09-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the disconnection between symbolic models, which lack complete audio recordings, and acoustic models, whose compositional logic remains implicit. To bridge this gap, we propose a framework that unifies music generation through symbolic planning. Methodologically, the approach employs an AR-NAR hybrid Transformer with a Mixture-of-Transformers (MoT) architecture, enabling a controllable β€œscore-first, audio-second” generation pipeline via semantic music tokenization. Furthermore, supervision mechanisms based on MERT2 and SheetSage2 are introduced to support zero-shot song cover generation. Experimental results demonstrate that the proposed model achieves a score of 6.73 on the WildSongBench benchmark, surpassing existing baselines. Additionally, expert preference evaluations indicate significant advantages over planning-free approaches and mainstream commercial systems, thereby realizing high-quality and interpretable music creation.
πŸ“ Abstract
Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through symbolic planning. A single AR-NAR Mixture-of-Transformers (MoT) first writes a readable score specifying melody and harmony, expands it into semantic music tokens, and realizes it as full-song audio. In comparisons using the same checkpoint, experts prefer symbolic planning for overall quality and musicality, with 49.3% of overall preferences versus 34.6% without planning. Experts also favor the unified model over a separate language model and diffusion Transformer. On WildSongBench, YuE2 scores 6.73 on SongBench Global Avg, exceeding all evaluated public baselines. Selecting from eight candidates (best-of-8), YuE2 reaches 6.96, the highest observed mean among all evaluated systems. Expert listening further establishes its competitiveness with proprietary song generators, favoring best-of-8 over Suno v4.5 and yielding nearly balanced preferences against Suno v5. To learn this generation process from recordings without aligned scores, we introduce MERT2 and SheetSage2 to supply semantic and symbolic supervision. MERT2 sets a new state of the art in music representation learning, surpassing previous best results on 14 of 15 MARBLE metrics; SheetSage2 leads 12 of 15 benchmark-metric pairs in our lead-sheet transcription comparison. The same checkpoint follows score edits while largely preserving unedited musical content and generates zero-shot covers without cover-specific training. Its readable score also enables agentic music editing, with external language models translating user feedback into revisions of the composition.
Problem

Research questions and friction points this paper is trying to address.

music generation
symbolic music
audio generation
unified model
controllability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Symbolic Planning
Mixture-of-Transformers
Unified Music Generation
Music Representation Learning
Agentic Music Editing