🤖 AI Summary
This work addresses critical challenges in voice cloning—namely, terminological inconsistency, lack of standardized evaluation criteria, and conceptual conflation across technical approaches—by establishing the first comprehensive, standardized taxonomy. It rigorously distinguishes two primary research paradigms: generative voice cloning (encompassing speaker adaptation, few-shot/zero-shot/multilingual TTS) and voice spoofing detection. Through a systematic survey of state-of-the-art methods from 2018 to 2024, it synthesizes core techniques—including deep neural architectures, self-supervised representations, meta-learning, and cross-lingual transfer—into a structured technical landscape. The paper also consolidates authoritative benchmark datasets and evaluation metrics, proposing a reproducible, unified evaluation protocol. Collectively, these contributions provide a foundational theoretical framework and practical guidelines for advancing voice cloning technologies, strengthening security governance, and informing ethical regulation.
📝 Abstract
Voice Cloning has rapidly advanced in today's digital world, with many researchers and corporations working to improve these algorithms for various applications. This article aims to establish a standardized terminology for voice cloning and explore its different variations. It will cover speaker adaptation as the fundamental concept and then delve deeper into topics such as few-shot, zero-shot, and multilingual TTS within that context. Finally, we will explore the evaluation metrics commonly used in voice cloning research and related datasets. This survey compiles the available voice cloning algorithms to encourage research toward its generation and detection to limit its misuse.