π€ AI Summary
This work addresses the limitations in current non-human voice conversion research, which has been hindered by the absence of publicly available, structured datasets and an overreliance on natural human speech. To bridge this gap, we present the first systematically constructed and open-sourced dataset encompassing both human and animal vocalizations, augmented with professionally designed audio effects. The dataset explicitly disentangles timbre and style dimensions and incorporates a structured train-test split strategy, enabling rigorous evaluation of modelsβ controllable generalization capabilities across both seen and unseen timbres and styles. By providing a high-quality benchmark and reproducible foundation, this resource significantly advances research in non-human voice conversion.
π Abstract
Advances in AI-based voice conversion have enabled a wide range of media applications, including films, audiobooks, and games. However, most research and public benchmarks still focus on natural human speech, leaving designed vocalizations, such as monster growls and robotic voices, underexplored, partly due to the lack of publicly available resources. To address this gap, we introduce the Designed Vocalizations Dataset, constructed by curating diverse raw vocal sources, including speech and animal vocalizations, and applying professional vocal effects processing to produce corresponding effect modified variants. We further provide a standardized test set with explicit seen/unseen splits over source timbre groups and preset styles to assess generalization under controlled conditions. Finally, we report baseline benchmark results to support reproducible evaluation and future research. The dataset and demo samples are available at https://ncai-official.github.io/speech/publications/designed-vocalizations-dataset/.