VibE-SVC: Vibrato Extraction with High-frequency F0 Contour for Singing Voice Conversion

📅 2025-05-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
In singing voice conversion, vibrato—characterized by periodic pitch modulation—is challenging to model explicitly and control accurately. This paper proposes an end-to-end controllable vibrato conversion framework to address this limitation. Our key innovation is the first application of discrete wavelet transform (DWT) to decompose the fundamental frequency (F0) contour into multi-scale components, explicitly isolating and representing the high-frequency dynamic modulations associated with vibrato—thereby overcoming the inherent opacity of conventional end-to-end models. Building upon this decomposition, we design a controllable architecture enabling fine-grained adjustment of vibrato attributes, including intensity and rate. Experiments demonstrate that our method preserves source singer identity while significantly improving the naturalness and expressiveness of vibrato transfer. Both objective metrics and subjective listening tests confirm consistent superiority over existing baseline approaches.

Technology Category

Natural Language Processing: SpeechComputer Vision: Diffusion Models for VisionSearch and Optimization: Mixed Discrete/Continuous Search

Application Category

User Modeling, Personalization and Recommendation: User modeling and simulation for interactive and conversational systemsGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsSearch and Retrieval-Augmented AI: Vertical and domain-specific search
📝 Abstract
Controlling singing style is crucial for achieving an expressive and natural singing voice. Among the various style factors, vibrato plays a key role in conveying emotions and enhancing musical depth. However, modeling vibrato remains challenging due to its dynamic nature, making it difficult to control in singing voice conversion. To address this, we propose VibESVC, a controllable singing voice conversion model that explicitly extracts and manipulates vibrato using discrete wavelet transform. Unlike previous methods that model vibrato implicitly, our approach decomposes the F0 contour into frequency components, enabling precise transfer. This allows vibrato control for enhanced flexibility. Experimental results show that VibE-SVC effectively transforms singing styles while preserving speaker similarity. Both subjective and objective evaluations confirm high-quality conversion.
Problem

Research questions and friction points this paper is trying to address.

Extracting and controlling vibrato in singing voice conversion
Modeling dynamic vibrato for expressive singing style transfer
Enhancing vibrato manipulation precision using discrete wavelet transform
Innovation

Methods, ideas, or system contributions that make the work stand out.

Explicit vibrato extraction using wavelet transform
Decomposes F0 contour for precise vibrato transfer
Controllable singing voice conversion model
🔎 Similar Papers
No similar papers found.
J
Joon-Seung Choi
Department of Artificial Intelligence, Korea University, Seoul, Korea
D
Dong-Min Byun
Department of Artificial Intelligence, Korea University, Seoul, Korea
Hyung-Seok Oh
Hyung-Seok Oh
Korea Unviersity
Speech synthesis Deep Learning
S
Seong-Whan Lee
Department of Artificial Intelligence, Korea University, Seoul, Korea