ARIA: Audio-Driven Melody-Tone Relation Modeling for Cantonese Lyric Authoring

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of aligning Cantonese lyrics with melody, where implicit pitch and noise interference in real-world singing audio hinder accurate tone-melody correspondence. To this end, we propose a two-stage generative framework. First, a tri-stream relation-aware tone estimator is designed to precisely predict tone sequences from timestamped audio. Subsequently, a decoupled retrieval-augmented generator, integrated with a Jyutping tone mapping technique, is constructed to produce prosodically compliant Cantonese lyrics. Additionally, a large-scale aligned dataset is established to support this novel task. Experimental results demonstrate that the proposed method achieves superior performance in both tone prediction and tone-consistent lyric generation, validating the effectiveness of multi-stream acoustic modeling and relational structure learning.
📝 Abstract
Cantonese lyric writing requires close alignment between lexical tones and melodic pitch. Existing melody-guided lyric generation methods typically rely on symbolic melody to generate lyrics. However, in real songwriting scenarios, melodies are often expressed as raw singing audio or hummed recordings, where pitch is implicit, noisy, and unstructured, making these methods difficult to apply directly. To address this limitation, we propose ARIA, a two-stage audio-driven melody-tone relation modeling framework for Cantonese lyric authoring that generates Cantonese lyrics from singing recordings with provided character-level timestamps. Specifically, we first design a Tri-Stream Relation-Aware Tone Estimator (TRATE) to predict 0243 sequences from timestamped singing audio by modeling multi-stream acoustic cues and relational tonal structure. We then propose a Decoupled Retrieval-Augmented Tone-Conditioned Lyric Generator (DRA-TCLG) to generate fluent lyrics conditioned on predicted tonal plans with retrieval-enhanced lexical guidance. Moreover, we construct a large-scale aligned audio-Jyutping-0243 dataset from real Cantonese singing recordings to support this new task. Experimental results demonstrate that ARIA achieves strong performance in both 0243 prediction and tone-consistent lyric generation, validating the effectiveness of the proposed framework.
Problem

Research questions and friction points this paper is trying to address.

Cantonese lyric generation
audio-driven
melody-tone alignment
singing recordings
Innovation

Methods, ideas, or system contributions that make the work stand out.

Audio-driven lyric generation
Melody-tone relation modeling
Tri-stream tone estimator
Retrieval-augmented generation
Cantonese singing dataset
🔎 Similar Papers