Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints

📅 2026-04-20
📈 Citations: 0
Influential: 0
📄 PDF

career value

195K/year
🤖 AI Summary
This work addresses the challenge that large language models often generate musically incoherent melodies from lyrics due to neglecting fundamental musical constraints, resulting in issues such as erratic rhythm and inappropriate pitch ranges. To mitigate this, the authors propose a novel two-stage alignment framework that requires no human annotations: first, preference data are automatically generated using unsupervised music-theoretic rules; then, domain knowledge is injected into the model by integrating Direct Preference Optimization (DPO) with Kahneman-Tversky Optimization (KTO). This approach uniquely unifies rule-based constraints with preference-based learning, significantly enhancing both the musical validity and perceptual quality of generated melodies. Experimental results demonstrate that the aligned model consistently outperforms existing baselines across objective metrics and subjective evaluations, producing outputs that are markedly more musical and coherent.

Technology Category

Application Category

📝 Abstract
Large Language Models (LLMs) show promise in lyric-to-melody generation, but models trained with Supervised Fine-Tuning (SFT) often produce musically implausible melodies with issues like poor rhythm and unsuitable vocal ranges, a phenomenon we term "constraint violation". To address this, we propose a novel alignment framework that instills musical knowledge without human annotation. We define rule-based musical constraints to automatically generate a preference dataset from an SFT model's outputs. The model is then aligned through a sequential process, first using Direct Preference Optimization (DPO) on paired preference data, followed by Kahneman-Tversky Optimization (KTO) on unpaired negative samples. Experimental results demonstrate that our aligned model substantially reduces rule violations and outperforms strong baselines in both objective and subjective evaluations, generating melodies with substantially improved musicality and coherence. An interactive demo with audio comparisons is available at https://arain233.github.io/AligningMelody-demo.
Problem

Research questions and friction points this paper is trying to address.

lyric-to-melody generation
constraint violation
musical plausibility
vocal range
rhythm
Innovation

Methods, ideas, or system contributions that make the work stand out.

rule-based musical constraints
preference dataset generation
Direct Preference Optimization (DPO)
Kahneman-Tversky Optimization (KTO)
lyric-to-melody generation
🔎 Similar Papers
H
Hao Meng
Zuoyebang Education Technology, Beijing, China
S
Siyuan Zheng
Zuoyebang Education Technology, Beijing, China
S
Shuran Zhou
Zuoyebang Education Technology, Beijing, China
Q
Qiangqiang Wang
Zuoyebang Education Technology, Beijing, China
Y
Yang Song
Zuoyebang Education Technology, Beijing, China