CLASVS: Continuous-Latent Autoregression for Melody-Preserving Lyric Editing in Singing Voice Synthesis

πŸ“… 2026-08-04
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the issue of generation bias arising from conflicts between source and target lyrics in continuous latent spaces by proposing a continuous latent autoregressive lyric editing method that requires neither discrete codebooks nor paired editing data. The approach introduces a State-Control-Transition routing mechanism coupled with a Progressive State-Control Grounding training strategy, enabling precise and flexible lyric replacement while preserving vocal rhythm, timbre, and naturalness. Experiments on two Mandarin singing datasets demonstrate that the proposed method substantially outperforms the discrete autoregressive model Vevo2 across four editing tasks, achieving a 46.2% reduction in macro phoneme error rate while effectively maintaining melodic accuracy, singer similarity, and perceptual quality.
πŸ“ Abstract
Reference-conditioned melody-preserving lyric editing replaces words while retaining a performance's timing, singer identity, and naturalness. Continuous-latent autoregression avoids finite codebooks and offers stepwise generation with learned stopping. Editing creates a conflict absent from ordinary reconstruction: training pairs reference cues with original lyrics, whereas inference asks revised lyrics to override source-lyric-correlated cues; one source-following patch can propagate through AR history. We introduce CLASVS. Its State-Control-Transition (SCT) routing keeps target-lyric and reference-melody controls persistent, returns semantic feedback on phonetic progress to the causal planner, and confines the previous latent patch to the local Transition. Progressive State-Control Grounding (PSCG) learns this contract through paired-edit-free, content-consistent Mandarin reconstruction. On two Mandarin benchmarks, CLASVS improves all four operations over discrete-AR Vevo2 and reduces macro-PER by 46.2%, while maintaining melody, singer similarity, and perceptual quality. Together, these results establish a strong continuous-AR operating point for score-annotation-free lyric edits and a basis for broader stepwise control. Audio demonstrations are available on our project page: https://piedpiperg.github.io/Liyric-SVS/.
Problem

Research questions and friction points this paper is trying to address.

lyric editing
singing voice synthesis
melody preservation
continuous-latent autoregression
reference-conditioned generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

continuous-latent autoregression
melody-preserving lyric editing
State-Control-Transition routing
singing voice synthesis
reference-conditioned generation
πŸ”Ž Similar Papers
No similar papers found.