A barrier or a booster? Familiarity effects on Mandarin emotion prosody recognition using AI-powered voice cloning

๐Ÿ“… 2026-09-29
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study investigates how acoustic anomalies in AI-synthesized speech and speaker familiarity interact to influence emotional prosody recognition and cognitive load. Employing a within-subjects design, the research integrates behavioral measures, including reaction time and accuracy, with physiological monitoring via heart rate variability (HRV). Notably, it is the first to combine subjective familiarity ratings with HRV to compare emotional decoding between human and AI-generated voices. The findings demonstrate that human speech yields higher recognition accuracy and faster responses, revealing that top-down sociocognitive mechanisms constrain the decoding of synthesized speech. Although HRV did not significantly differentiate experimental conditions, the results underscore current limitations in AIโ€™s capacity to replicate emotional prosody. Ultimately, this work provides empirical evidence to inform the optimization of humanโ€“machine voice interaction systems.
๐Ÿ“ Abstract
Emotion prosody perception requires simultaneous processing of acoustic cues and speaker identity. While listeners effortlessly decode natural speech, AI synthetic voices introduce cognitive complexities due to subtle acoustic atypicalities. It remains unclear how these synthetic features interact with a listener's prior social knowledge and memory of a familiar speaker. This study investigated how speech sources (human vs. AI) and speaker familiarity affect emotion recognition accuracy and cognitive load. A within-subject task with Mandarin-speaking adults evaluated behavioral (accuracy, reaction time) and physiological data (heart rate variability). Results showed that human voices yielded significantly higher accuracy and faster processing times than AI voices, while HRV did not significantly differentiate between conditions. These findings show that decoding synthetic speech is gated by top-down social cognition, highlighting limitations in current AI synthesis technologies.
Problem

Research questions and friction points this paper is trying to address.

emotion prosody recognition
AI voice cloning
speaker familiarity
cognitive load
synthetic speech
Innovation

Methods, ideas, or system contributions that make the work stand out.

AI voice cloning
emotion prosody recognition
speaker familiarity
heart rate variability
cognitive load
๐Ÿ”Ž Similar Papers
F
Feng Xu
Institute of Software, Chinese Academy of Sciences
G
Gaoyuan Zhang
Institute of Software, Chinese Academy of Sciences
S
Shanshan Xue
Institute of Software, Chinese Academy of Sciences
Y
Yixiang Chen
Institute of Software, Chinese Academy of Sciences
H
Hanrui Zhou
Institute of Software, Chinese Academy of Sciences
Xurong Xie
Xurong Xie
Institute of Software, Chinese Academy of Sciences
Speech and Language ProcessingMachine LearningHuman-Computer InteractionAI for Health
H
Hui Chen
Institute of Software, Chinese Academy of Sciences