🤖 AI Summary
This study addresses the quality validation challenge in constructing AI voice cloning corpora by evaluating the perceptual equivalence of synthetic and natural speech in intonation perception tasks. Using a Singing Voice Conversion (SVC) model to generate synthetic stimuli, we designed a dual-experiment paradigm for psycholinguistic behavioral testing, systematically comparing similarity ratings and intonation recognition performance across varying familiarity levels. Results reveal that interrogative sentences, serving as speaker identification cues, are particularly susceptible to interference from synthetic artifacts. Significant interaction effects were observed among speech type, intonation, and familiarity, elucidating how speech type modulates the processing of familiarity. These findings confirm the critical roles of intonational features and familiarity in synthetic speech perception, offering empirical guidance for improving corpus construction standards in voice cloning research.
📝 Abstract
Language training relies on a corpus constructed by a large number linguistic materials. AI-powered voice clones provide a way to construct the corpus with relatively low cost. Singing voice conversion (SVC) model is used to generate synthetic voices. This study compares participants' performances on natural and synthetic speech in two experiments, similarity perception and intonation recognition. In the accuracy of similarity perception task, a significant interaction between speech type and intonation is found, suggesting that question may serve as a cue for speaker identification but may be influenced by synthetic features. In the accuracy of intonation recognition task, a significant interaction between speech type and familiarity is observed, indicating that speech type affects how much familiarity contributes to voice processing.