๐ค AI Summary
This study investigates how acoustic anomalies in AI-synthesized speech and speaker familiarity interact to influence emotional prosody recognition and cognitive load. Employing a within-subjects design, the research integrates behavioral measures, including reaction time and accuracy, with physiological monitoring via heart rate variability (HRV). Notably, it is the first to combine subjective familiarity ratings with HRV to compare emotional decoding between human and AI-generated voices. The findings demonstrate that human speech yields higher recognition accuracy and faster responses, revealing that top-down sociocognitive mechanisms constrain the decoding of synthesized speech. Although HRV did not significantly differentiate experimental conditions, the results underscore current limitations in AIโs capacity to replicate emotional prosody. Ultimately, this work provides empirical evidence to inform the optimization of humanโmachine voice interaction systems.
๐ Abstract
Emotion prosody perception requires simultaneous processing of acoustic cues and speaker identity. While listeners effortlessly decode natural speech, AI synthetic voices introduce cognitive complexities due to subtle acoustic atypicalities. It remains unclear how these synthetic features interact with a listener's prior social knowledge and memory of a familiar speaker. This study investigated how speech sources (human vs. AI) and speaker familiarity affect emotion recognition accuracy and cognitive load. A within-subject task with Mandarin-speaking adults evaluated behavioral (accuracy, reaction time) and physiological data (heart rate variability). Results showed that human voices yielded significantly higher accuracy and faster processing times than AI voices, while HRV did not significantly differentiate between conditions. These findings show that decoding synthetic speech is gated by top-down social cognition, highlighting limitations in current AI synthesis technologies.