hierarchical music modeling

Designs and implements models, representation schemes, and neural architectures that capture musical structure at multiple hierarchical levels—using hierarchical latent variables, dual-stage pipelines, and hierarchical representation learning—to enable both global planning and local refinement. Builds generation and inference methods that operate on mixed-token and track-specific representations (e.g., separate vocal and accompaniment tokens), allowing prediction of high-level semantic tokens and subsequent refinement of track-specific acoustic detail while preserving overall musical coherence.

hierarchicalmusicmodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.64
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing language models struggle to simultaneously maintain global coherence, musicality, vocal and accompaniment fidelity, and lyrical adherence in full-length song generation. This work proposes a hybrid LLM-diffusion framework that first employs a hierarchical language model (LeLM) for semantic planning, then parallelizes the refinement of vocals and accompaniment, and finally reconstructs high-quality waveforms using a diffusion-based music codec. A novel aesthetics-guided progressive post-training strategy decouples musicality learning, controllability alignment, and acoustic refinement, supported by an automated music aesthetics evaluator during pretraining. Combining supervised fine-tuning (SFT), large-scale offline DPO, and closed-loop semi-online DPO, the system significantly outperforms open-source baselines across six subjective dimensions and approaches leading commercial systems on multiple perceptual metrics, with ablation studies confirming the efficacy of both architecture and training methodology.

coherence and musicalityfull-length song generationlyrics and prompt alignment

TOMI: Transforming and Organizing Music Ideas for Multi-Track Compositions with Full-Song Structure

Jun 29, 2025
QH
Qi He
🏛️ Music X Lab | MBZUAI | New York University

To address limitations in creative generation, structural organization, and human-AI collaboration in multi-track electronic music composition, this paper introduces TOMI—a novel framework that pioneers a sparse four-dimensional representation integrating conceptual hierarchy and spatiotemporal structure (segment–section–track–transformation operation), enabling end-to-end composition via instruction-tuned large language models. TOMI unifies MIDI/audio generation and conversion techniques to synthesize full-song multi-track arrangements and natively integrates with the REAPER digital audio workstation. Experimental evaluations demonstrate that TOMI significantly outperforms baseline methods in musical structural coherence and creative expression quality. A user study further confirms its effectiveness in supporting complete multi-track song generation and efficient human-AI co-creation within authentic production workflows.

Enabling interactive human-AI co-creation in music productionGenerating multi-track music with full-song structural coherenceModeling hierarchical music composition across time and space

Learning Relationships Between Separate Audio Tracks for Creative Applications

Sep 29, 2025
BB
Balthazar Bujard
🏛️ STMS Lab Ircam | CNRS | Sorbonne Université

This work addresses the challenge of modeling musical relationships between input and generated output in real-time music interaction. We propose a music agent training framework leveraging a separated-track database. Methodologically, we introduce the first end-to-end architecture integrating a symbolic decision-making module: a Transformer models and predicts musical relationships symbolically; Wav2Vec 2.0 serves as the perception module for audio representation extraction; and concatenative synthesis enables high-fidelity audio rendering. Our key contribution is the explicit encoding of pairwise track co-relationships—e.g., A→B—as a learnable symbolic decision process, a novel formulation in generative music systems. Experiments demonstrate that the model accurately reproduces the musical relationships observed in training data (e.g., A→B mappings), and under real-time guidance, generates semantically coherent and stylistically consistent response tracks. This significantly enhances controllability and expressive capability in creative music applications.

Learning musical relationships between paired audio tracks for creative applicationsReal-time generation of coherent musical output guided by live inputTraining symbolic decision modules to predict musical relationships from separated tracks

This work addresses the limitations of existing symbolic music tokenization approaches, which typically rely on event sequences with irregular time steps and struggle to explicitly model rhythmic regularities. The authors propose a novel tokenization method that uses fixed temporal units—such as beats—as the fundamental building blocks, merging all events sharing the same pitch within a single time step into one token, thereby yielding a sparse piano-roll-like representation. This approach enables explicit alignment of temporal structure for the first time. Integrated with a Transformer architecture, it significantly enhances generation quality, structural coherence, and long-range dependency modeling in tasks such as music continuation and accompaniment generation, while also achieving superior efficiency and rhythmic consistency compared to prevailing event-based tokenization methods.

language modelsmusical time representationsymbolic music

Existing song generation systems often suffer from abrupt section transitions, insufficient dynamic range, and monotonous arrangements due to the absence of explicit structural planning and fine-grained multi-track modeling. This work proposes a hierarchical generative framework that first predicts high-level sketch tokens from compressed audio representations to outline the global song structure, then conditions fine-grained audio generation on this sketch. The approach explicitly models four distinct tracks—vocals, bass, drums, and other instruments—to accurately capture the role and interaction of each musical part. By integrating hierarchical architecture, sketch-guided generation, and parallel multi-track modeling, the method substantially enhances structural coherence and arrangement richness. It outperforms baseline models on both objective metrics and human listening tests, achieving generation quality comparable to strong open-source systems—even without lyric alignment or preference-based optimization.

arrangement planningcoherencemulti-track modeling

Latest Papers

What's happening recently
View more

This work proposes a unified framework for full-song generation that simultaneously supports lyric-conditioned singing, instrumental music, and cover song synthesis. The approach integrates hierarchical autoregressive modeling with continuous flow-matching rendering and introduces a dual-level melody guidance mechanism to preserve the original melodic structure in cover songs. For the first time, reward-driven optimization strategies—including DPO, GRPO, and OPD—are applied to full-song generation. The system employs a semantic-aware 8-codebook RVQ audio tokenizer, a Hybrid-LM language model, and a FullDiT decoder. Evaluations on multilingual automatic metrics and the Artificial Analysis Music with Vocals leaderboard demonstrate state-of-the-art performance in song completeness, audio fidelity, and musicality.

cover song generationfull-song generationinstrumental music generation

This work addresses the limitations of existing collaborative music-generation agents, which lack internal representations that jointly support understanding and generation while preserving human creative control. We propose a hierarchical self-supervised world model trained on unlabeled MIDI piano rolls using a Swin V2 encoder and a JEPA-style objective to learn musical structure. High-quality generation and interactive prompting are achieved via conditional flow matching. Notably, our approach spontaneously discovers temporal and phrase-level structures without prior music-theoretic knowledge, enhances interpretability through hierarchical embeddings, and enables graphical mask-based inpainting without specialized samplers. Experiments show that frozen embeddings disentangle musical attributes across timescales; adding a lightweight chord-supervision head boosts chord recovery accuracy from 0.18 to 0.54 and achieves a tonality detection score of 0.70. The model attains a generation F1 score of 0.996, with inference times of 2.8 seconds on CPU and just 0.6 seconds on Apple MPS, and has been integrated into a real-time interactive system.

human-AI collaborationmusic co-creationself-supervised learning

Existing neural network models struggle to generate cohesive, memorable game music with consistent repetition due to their limited understanding of musical structure, hindering their practical application in gaming contexts. To address this limitation, this work introduces a novel dataset comprising 309 structurally annotated game music audio tracks and proposes the first supervised learning approach for music structure segmentation in this domain. By employing a hybrid CNN-RNN architecture, the model achieves segmentation performance on par with current state-of-the-art unsupervised methods despite using significantly less training data. These results demonstrate the effectiveness and potential of supervised learning for modeling the structural characteristics of game music, offering a promising direction for future research in controllable and structure-aware music generation.

game musicmusic segmentationmusical structure

This work addresses the limited understanding of how current audio diffusion models internally represent high-level musical semantics—such as instruments, vocals, style, rhythm, and emotion—and the consequent lack of precise control over generation. By applying activation patching to identify critical attention layers and combining contrastive activation addition with sparse autoencoders, we reveal—for the first time—the existence of shared yet functionally specialized subsets of attention heads that explicitly encode distinct high-level musical concepts. Leveraging this insight, we demonstrate high-precision intervention and editing of specific musical elements in generated audio, substantially enhancing both the interpretability and controllability of diffusion-based audio synthesis.

activation steeringaudio diffusion modelshigh-level representation

This work addresses the complexity and optimization challenges of traditional high-fidelity music generation, which relies on heterogeneous representations separating structure from detail. The authors propose a unified two-stage generation framework within a single deep acoustic token space: a backbone model first generates coarse-grained tokens for the full composition, followed by a super-resolution model that progressively refines these tokens in parallel layers within the same space. They demonstrate, for the first time, that a purely acoustic token-based language model can spontaneously align lyrics with vocals without requiring a separate semantic stage. Initializing the super-resolution model with the backbone significantly accelerates convergence and enhances audio quality. Leveraging a 64-layer RVQ representation, hybrid attention (causal for alignment, full for refinement), and fixed 62-step efficient inference, the method achieves high-fidelity audio reconstruction while preserving precise lyric-vocal alignment, validating the efficacy and superiority of the unified acoustic token paradigm.

acoustic tokenhigh-fidelitymusic generation

Hot Scholars

JH

Junjun He

Shanghai Jiao Tong University
LL

Liang Lin

Fellow of IEEE/IAPR, Professor of Computer Science, Sun Yat-sen University
Embodied AICausal Inference and LearningMultimodal Data Analysis
LC

Longbing Cao

Distinguished Chair Professor in AI & ARC Future Fellow (Level 3), Macquarie University
Artificial intelligenceData scienceMachine learningBehavior informatics
DS

Daniel Shao

PhD Candidate, Massachusetts Institute of Technology
machine learningfoundation modelscomputer vision