Rubric-Based Optimization for Text-to-Music Generation

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of post-training for text-to-music generation, which typically relies on single reward metrics that fail to comprehensively capture multidimensional musical quality. To overcome this, the authors propose a structured reward framework grounded in pretrained audio language models (ALMs). A key innovation is a division-of-labor strategy pairing ALM-based scoring rubrics with dedicated objective rewards: the former evaluates perceptual qualities that resist formalization, while the latter precisely optimizes quantifiable attributes such as rhythm and tonality. This framework is integrated with DPO and DiffusionNFT algorithms to align autoregressive and diffusion-based generative models, respectively. Experiments on the MusicCaps benchmark demonstrate simultaneous improvements across multiple metrics, including CLAP and SongEval, validating the effectiveness of structured rewards in synergistically optimizing multidimensional musical quality.
📝 Abstract
Post-training text-to-music generation requires reward signals that capture multiple aspects of musical quality beyond what any single automatic metric can measure. We study structured, rubric-based rewards from pretrained audio-language models (ALMs) as training signals for both autoregressive and diffusion-based music generators. An ALM scores each generated clip against the rubric; we rank candidates generated for the same text prompt by their scores and convert these rankings into preference pairs for DPO on both MusicGen-small and ACE-Step~v1, and additionally use the rubric scores directly as scalar rewards for DiffusionNFT on ACE-Step~v1. On MusicCaps, rubric-based optimization improves CLAP, SongEval, and Audiobox-Aesthetics simultaneously, with the strongest gains obtained by DiffusionNFT on ACE-Step. By contrast, on MusicGen-small, building preferences from any one of these automatic evaluators produces clear cross-metric trade-offs: the targeted evaluator improves while other independent evaluators deteriorate. We further study tempo, key, and instrumentation, where precise objective rewards are available. Directly optimizing these specialized rewards reliably improves the target attributes, whereas ALM rubrics provide only partial transfer for tempo and instrumentation and no measurable improvement for key. Together, these results suggest a practical division of labor: ALM rubrics are effective for broad perceptual qualities that are difficult to formalize, while specialized objective rewards remain preferable when reliable measurements are available.
Problem

Research questions and friction points this paper is trying to address.

text-to-music generation
reward signal
rubric-based evaluation
audio-language models
post-training optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Rubric-Based Optimization
Text-to-Music Generation
Audio-Language Models
Direct Preference Optimization
DiffusionNFT
🔎 Similar Papers
2024-09-30Citations: 2