Score
Designs, implements, and evaluates audio systems that transform or reproduce human vocal identity and style — for example, models that convert a source speaker’s speech to sound like a target speaker, clone a specific voice from limited data, or synthesize singing in a target voice — by modeling and controlling timbre, pitch, prosody, timing, and linguistic content. Develops and analyzes the algorithms, training procedures, and adaptation techniques (including few‑shot or zero‑shot cloning) used to preserve intelligibility and expressiveness while changing or generating speaker identity.
This work addresses critical challenges in voice cloning—namely, terminological inconsistency, lack of standardized evaluation criteria, and conceptual conflation across technical approaches—by establishing the first comprehensive, standardized taxonomy. It rigorously distinguishes two primary research paradigms: generative voice cloning (encompassing speaker adaptation, few-shot/zero-shot/multilingual TTS) and voice spoofing detection. Through a systematic survey of state-of-the-art methods from 2018 to 2024, it synthesizes core techniques—including deep neural architectures, self-supervised representations, meta-learning, and cross-lingual transfer—into a structured technical landscape. The paper also consolidates authoritative benchmark datasets and evaluation metrics, proposing a reproducible, unified evaluation protocol. Collectively, these contributions provide a foundational theoretical framework and practical guidelines for advancing voice cloning technologies, strengthening security governance, and informing ethical regulation.
This study challenges the common misconception that mainstream voice cloning technologies faithfully replicate a speaker’s voice, demonstrating instead that they systematically introduce stylistic biases. Through comprehensive human subjective evaluations, acoustic feature analyses (e.g., accent and speech rate), and variance measurements in audio embedding spaces, the work reveals that cloned voices are consistently perceived as more authoritative, warmer, customer-service-like, and anthropomorphic. These perceptual shifts significantly increase user trust and willingness to disclose sensitive information. Concurrently, the diversity of vocal characteristics is markedly reduced, leading to voice homogenization that threatens speaker identity authenticity and compromises user behavioral security. This paper thus establishes that current voice cloning operates fundamentally as a form of style transfer rather than accurate vocal reproduction.
Addressing key challenges in singing voice synthesis (SVS) and singing voice conversion (SVC)—including cross-domain modeling difficulty, insufficient musicality, and scarcity of high-quality annotated data—this paper proposes the first voice-reference-driven zero-shot unified framework for SVS/SVC. Methodologically, it employs a pretrained content encoder to extract shared phonetic–singing representations, integrates a diffusion-based generative model trained jointly on hybrid singing/speech data, and introduces a multi-condition controllable decoding mechanism. This enables fully controllable generation of lyrics, pitch, style, and timbre from a single spoken utterance alone. Experiments demonstrate significant improvements over state-of-the-art methods in timbre similarity and musicality; notably, the framework achieves high-fidelity singing voice cloning under zero-shot conditions. By eliminating reliance on target-domain singing data or parallel annotations, it establishes a novel paradigm for low-resource music generation.
Existing text-to-speech (TTS) models struggle to directly modify reference timbres and achieve segment-level local expressiveness control. To address these limitations, this work proposes EDICT, a framework that pioneers the construction of shared timbre anchors within the codec token space to unify global timbre editing with local expressiveness control. Specifically, the method generates edited acoustic references to anchor target timbres and integrates segment-wise instructions with dynamic KV cache reconstruction techniques, effectively balancing instruction following, speaker consistency, and transition quality. Experimental results demonstrate that EDICT significantly improves timbre editing performance and overall quality on benchmarks such as TimbreEdit-Bench.
Existing singing voice synthesis (SVS) methods lack flexible, natural-language-based control over stylistic attributes such as gender, vocal range, and loudness. This work introduces the first text-driven controllable SVS system enabling explicit, fine-grained multi-attribute editing. Methodologically: (1) we propose a multi-scale pitch representation that disentangles vocal range from melody; (2) we design a text encoder fine-tuning strategy coupled with cross-modal (speech + singing) data augmentation; and (3) we pioneer end-to-end integration of natural language instructions into the SVS generation pipeline using a decoder-only transformer architecture. Experiments demonstrate that our system significantly outperforms baselines in both control accuracy and audio naturalness, while achieving high melodic fidelity and superior sound quality. All generated audio samples are publicly released for reproducibility and verification.
This paper addresses out-of-domain (OOD) zero-shot singing voice style transfer—specifically, transferring singing styles from unseen reference vocals encompassing timbre, emotion, articulation, and vocal technique. Methodologically, we propose the first end-to-end framework featuring: (i) a Residual Style Adapter (RSA) that explicitly models multi-dimensional singing style; (ii) Uncertainty-aware Modulated Layer Normalization (UMLN) to enhance cross-domain generalization; and (iii) an integrated design combining residual vector quantization, reference-driven disentangled style encoding, and neural acoustic modeling. Experiments demonstrate substantial improvements over state-of-the-art baselines in zero-shot transfer: MOS scores increase by over 1.2, cosine similarity for style fidelity improves by 23%, and both audio naturalness and expressiveness are significantly enhanced. Crucially, our approach eliminates reliance on target attributes observed during training—a fundamental limitation of conventional singing voice synthesis (SVS) systems.
Existing diffusion models struggle to unify voice and singing voice conversion with limited generalization capabilities. This work proposes the first adaptation of a multi-instrument music synthesis diffusion model to vocal conversion tasks, leveraging phonetic posteriorgrams (PPGs) and pitch contours as conditioning signals and incorporating feature-wise linear modulation (FiLM) to model speaker/singer identity within a unified speech–singing conversion framework. The approach requires no manual annotations and enables large-scale self-supervised training using off-the-shelf feature extractors. It achieves naturalness and performer similarity on par with or surpassing specialized systems while maintaining precise pitch control, thereby demonstrating the feasibility of cross-domain model transfer. However, it exhibits limitations in phonetic fidelity and experiences audio quality degradation due to the inclusion of instrumental data during training.
This study addresses the quality validation challenge in constructing AI voice cloning corpora by evaluating the perceptual equivalence of synthetic and natural speech in intonation perception tasks. Using a Singing Voice Conversion (SVC) model to generate synthetic stimuli, we designed a dual-experiment paradigm for psycholinguistic behavioral testing, systematically comparing similarity ratings and intonation recognition performance across varying familiarity levels. Results reveal that interrogative sentences, serving as speaker identification cues, are particularly susceptible to interference from synthetic artifacts. Significant interaction effects were observed among speech type, intonation, and familiarity, elucidating how speech type modulates the processing of familiarity. These findings confirm the critical roles of intonational features and familiarity in synthetic speech perception, offering empirical guidance for improving corpus construction standards in voice cloning research.
This study addresses the misuse of AI-generated speech and associated detection challenges by presenting the first systematic integration of generation and detection technologies across the full pipeline. Through constructing a technical taxonomy, curating benchmark resources, and establishing an open challenge framework, this work develops a comprehensive knowledge map of the field. The research not only clarifies technological evolution and critical bottlenecks but also delineates a future roadmap tailored to speech-specific characteristics. By providing systematic theoretical support and practical guidance for building robust speech security defenses, this survey fills a significant gap in existing literature regarding holistic, end-to-end perspectives on AI speech synthesis and forensics.
This work proposes a few-shot voice cloning system to address the scarcity of multi-speaker speech synthesis for low-resource Nepali. Leveraging a small amount of untranscribed Nepali speech, the system trains a speaker encoder and integrates it with a Tacotron2 acoustic model and a WaveRNN vocoder to generate target-speaker utterances directly from Devanagari script. A novel generative end-to-end loss is introduced to optimize speaker embeddings, and their representational quality is validated through UMAP visualization. Experimental results demonstrate that the system not only effectively clones voices of seen speakers but also generalizes to unseen speakers, achieving the first successful implementation of multi-speaker voice cloning for Nepali under low-resource conditions and offering a viable pathway toward personalized text-to-speech synthesis for other resource-constrained languages.
Existing song generation methods lack zero-shot speaker cloning capabilities, while voice conversion approaches typically neglect the joint generation of vocals and accompaniment. To address these limitations, this work proposes UniSinger—the first end-to-end unified framework that simultaneously supports accompaniment-aware speaker-cloned singing voice generation and singing voice conversion. Built upon a multimodal diffusion Transformer, UniSinger constructs a unified speaker embedding space and introduces task-specific modality masking alongside a curriculum learning strategy to harmonize multi-task optimization and mitigate task interference. The model achieves state-of-the-art performance on both tasks, enabling for the first time cross-task timbre control and mutual performance gains, thereby advancing the frontier of intelligent music generation.