voice conversion

Designs, implements, and evaluates audio systems that transform or reproduce human vocal identity and style — for example, models that convert a source speaker’s speech to sound like a target speaker, clone a specific voice from limited data, or synthesize singing in a target voice — by modeling and controlling timbre, pitch, prosody, timing, and linguistic content. Develops and analyzes the algorithms, training procedures, and adaptation techniques (including few‑shot or zero‑shot cloning) used to preserve intelligibility and expressiveness while changing or generating speaker identity.

voiceconversion

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.3
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$167K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study challenges the common misconception that mainstream voice cloning technologies faithfully replicate a speaker’s voice, demonstrating instead that they systematically introduce stylistic biases. Through comprehensive human subjective evaluations, acoustic feature analyses (e.g., accent and speech rate), and variance measurements in audio embedding spaces, the work reveals that cloned voices are consistently perceived as more authoritative, warmer, customer-service-like, and anthropomorphic. These perceptual shifts significantly increase user trust and willingness to disclose sensitive information. Concurrently, the diversity of vocal characteristics is markedly reduced, leading to voice homogenization that threatens speaker identity authenticity and compromises user behavioral security. This paper thus establishes that current voice cloning operates fundamentally as a form of style transfer rather than accurate vocal reproduction.

human perceptionspeaker homogenizationspeech synthesis

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference

Jan 23, 2025
SD
Shuqi Dai
🏛️ Carnegie Mellon University | Princeton University | Adobe Research

Addressing key challenges in singing voice synthesis (SVS) and singing voice conversion (SVC)—including cross-domain modeling difficulty, insufficient musicality, and scarcity of high-quality annotated data—this paper proposes the first voice-reference-driven zero-shot unified framework for SVS/SVC. Methodologically, it employs a pretrained content encoder to extract shared phonetic–singing representations, integrates a diffusion-based generative model trained jointly on hybrid singing/speech data, and introduces a multi-condition controllable decoding mechanism. This enables fully controllable generation of lyrics, pitch, style, and timbre from a single spoken utterance alone. Experiments demonstrate significant improvements over state-of-the-art methods in timbre similarity and musicality; notably, the framework achieves high-fidelity singing voice cloning under zero-shot conditions. By eliminating reliance on target-domain singing data or parallel annotations, it establishes a novel paradigm for low-resource music generation.

Domain AdaptationLimited Singing DataVoice Synthesis

Existing text-to-speech (TTS) models struggle to directly modify reference timbres and achieve segment-level local expressiveness control. To address these limitations, this work proposes EDICT, a framework that pioneers the construction of shared timbre anchors within the codec token space to unify global timbre editing with local expressiveness control. Specifically, the method generates edited acoustic references to anchor target timbres and integrates segment-wise instructions with dynamic KV cache reconstruction techniques, effectively balancing instruction following, speaker consistency, and transition quality. Experimental results demonstrate that EDICT significantly improves timbre editing performance and overall quality on benchmarks such as TimbreEdit-Bench.

expressive speech synthesislocal instruction controltext-to-speech

Existing singing voice synthesis (SVS) methods lack flexible, natural-language-based control over stylistic attributes such as gender, vocal range, and loudness. This work introduces the first text-driven controllable SVS system enabling explicit, fine-grained multi-attribute editing. Methodologically: (1) we propose a multi-scale pitch representation that disentangles vocal range from melody; (2) we design a text encoder fine-tuning strategy coupled with cross-modal (speech + singing) data augmentation; and (3) we pioneer end-to-end integration of natural language instructions into the SVS generation pipeline using a decoder-only transformer architecture. Experiments demonstrate that our system significantly outperforms baselines in both control accuracy and audio naturalness, while achieving high melodic fidelity and superior sound quality. All generated audio samples are publicly released for reproducibility and verification.

Singing Voice SynthesisStyle TransferVoice Conversion

StyleSinger: Style Transfer for Out-of-Domain Singing Voice Synthesis

Dec 17, 2023
YZ
Yu Zhang
🏛️ Zhejiang University | Huawei Cloud

This paper addresses out-of-domain (OOD) zero-shot singing voice style transfer—specifically, transferring singing styles from unseen reference vocals encompassing timbre, emotion, articulation, and vocal technique. Methodologically, we propose the first end-to-end framework featuring: (i) a Residual Style Adapter (RSA) that explicitly models multi-dimensional singing style; (ii) Uncertainty-aware Modulated Layer Normalization (UMLN) to enhance cross-domain generalization; and (iii) an integrated design combining residual vector quantization, reference-driven disentangled style encoding, and neural acoustic modeling. Experiments demonstrate substantial improvements over state-of-the-art baselines in zero-shot transfer: MOS scores increase by over 1.2, cosine similarity for style fidelity improves by 23%, and both audio naturalness and expressiveness are significantly enhanced. Crucially, our approach eliminates reliance on target attributes observed during training—a fundamental limitation of conventional singing voice synthesis (SVS) systems.

Overcoming out-of-domain quality declineStyle transfer for singing synthesisZero-shot style transfer enhancement

Latest Papers

What's happening recently
View more

Existing diffusion models struggle to unify voice and singing voice conversion with limited generalization capabilities. This work proposes the first adaptation of a multi-instrument music synthesis diffusion model to vocal conversion tasks, leveraging phonetic posteriorgrams (PPGs) and pitch contours as conditioning signals and incorporating feature-wise linear modulation (FiLM) to model speaker/singer identity within a unified speech–singing conversion framework. The approach requires no manual annotations and enables large-scale self-supervised training using off-the-shelf feature extractors. It achieves naturalness and performer similarity on par with or surpassing specialized systems while maintaining precise pitch control, thereby demonstrating the feasibility of cross-domain model transfer. However, it exhibits limitations in phonetic fidelity and experiences audio quality degradation due to the inclusion of instrumental data during training.

audio generationcross-domain adaptationdiffusion models

This study addresses the quality validation challenge in constructing AI voice cloning corpora by evaluating the perceptual equivalence of synthetic and natural speech in intonation perception tasks. Using a Singing Voice Conversion (SVC) model to generate synthetic stimuli, we designed a dual-experiment paradigm for psycholinguistic behavioral testing, systematically comparing similarity ratings and intonation recognition performance across varying familiarity levels. Results reveal that interrogative sentences, serving as speaker identification cues, are particularly susceptible to interference from synthetic artifacts. Significant interaction effects were observed among speech type, intonation, and familiarity, elucidating how speech type modulates the processing of familiarity. These findings confirm the critical roles of intonational features and familiarity in synthetic speech perception, offering empirical guidance for improving corpus construction standards in voice cloning research.

Equivalence AssessmentIntonation PerceptionSinging Voice Conversion

This study addresses the misuse of AI-generated speech and associated detection challenges by presenting the first systematic integration of generation and detection technologies across the full pipeline. Through constructing a technical taxonomy, curating benchmark resources, and establishing an open challenge framework, this work develops a comprehensive knowledge map of the field. The research not only clarifies technological evolution and critical bottlenecks but also delineates a future roadmap tailored to speech-specific characteristics. By providing systematic theoretical support and practical guidance for building robust speech security defenses, this survey fills a significant gap in existing literature regarding holistic, end-to-end perspectives on AI speech synthesis and forensics.

AI-generated voicedeepfake audiodisinformation

This work proposes a few-shot voice cloning system to address the scarcity of multi-speaker speech synthesis for low-resource Nepali. Leveraging a small amount of untranscribed Nepali speech, the system trains a speaker encoder and integrates it with a Tacotron2 acoustic model and a WaveRNN vocoder to generate target-speaker utterances directly from Devanagari script. A novel generative end-to-end loss is introduced to optimize speaker embeddings, and their representational quality is validated through UMAP visualization. Experimental results demonstrate that the system not only effectively clones voices of seen speakers but also generalizes to unseen speakers, achieving the first successful implementation of multi-speaker voice cloning for Nepali under low-resource conditions and offering a viable pathway toward personalized text-to-speech synthesis for other resource-constrained languages.

few-shotlow-resourcemulti-speaker

Existing song generation methods lack zero-shot speaker cloning capabilities, while voice conversion approaches typically neglect the joint generation of vocals and accompaniment. To address these limitations, this work proposes UniSinger—the first end-to-end unified framework that simultaneously supports accompaniment-aware speaker-cloned singing voice generation and singing voice conversion. Built upon a multimodal diffusion Transformer, UniSinger constructs a unified speaker embedding space and introduces task-specific modality masking alongside a curriculum learning strategy to harmonize multi-task optimization and mitigate task interference. The model achieves state-of-the-art performance on both tasks, enabling for the first time cross-task timbre control and mutual performance gains, thereby advancing the frontier of intelligent music generation.

accompaniment co-generationsinging voice conversionsong generation

Hot Scholars

ÉS

Éva Székely

Assistant Professor, KTH Royal Institute of Technology
speech technologyspeech synthesisdeep learninggenerative modelling
BS

Berrak Sisman

Assistant Professor (ECE & DSAI), Johns Hopkins University
Machine LearningAffective ComputingSpeech SynthesisVoice Conversion
HL

Haizhou Li

The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), China; NUS, Singapore
Automatic Speech RecognitionSpeaker RecognitionLanguage RecognitionVoice Conversion
XM

Xiaoxiao Miao

Duke Kunshan University
Speech PrivacySpeaker and Language IdentificationSpeech Synthesis