🤖 AI Summary
This study addresses the misuse of AI-generated speech and associated detection challenges by presenting the first systematic integration of generation and detection technologies across the full pipeline. Through constructing a technical taxonomy, curating benchmark resources, and establishing an open challenge framework, this work develops a comprehensive knowledge map of the field. The research not only clarifies technological evolution and critical bottlenecks but also delineates a future roadmap tailored to speech-specific characteristics. By providing systematic theoretical support and practical guidance for building robust speech security defenses, this survey fills a significant gap in existing literature regarding holistic, end-to-end perspectives on AI speech synthesis and forensics.
📝 Abstract
The ability of artificial intelligence (AI) models to generate highly realistic human voices has advanced rapidly. These technologies power accessibility tools, virtual assistants and creative applications, but they also enable harmful uses, including impersonation, fraud and disinformation. Recent incidents of voice cloning scams targeting businesses and political leaders underscore the urgent need for robust safeguards. Unlike image and video deepfakes, the detection of synthetic voices poses unique challenges due to the complexity of phonetics, prosody and auditory perception. This survey offers a comprehensive overview of AI voice generation and detection methods, encompassing both the technical foundations and the latest state-of-the-art advances. This study also identifies key open challenges, benchmark resources and future directions to make this survey useful for future researchers.