🤖 AI Summary
Commercial text-to-speech systems often reinforce gender stereotypes and power inequities through gendered and sexualized vocal qualities. Drawing on feminist human-computer interaction frameworks, this study systematically investigates listeners’ perceptions of male- and female-coded synthetic voices across neutral and sexualized scripts, employing auditory experiments, adjective rating scales, open-ended textual feedback, and acoustic analysis. The research reveals, for the first time, that mainstream AI voices exhibit a strongly binary and heteronormative construction of gender: female-coded voices are consistently perceived as more sexualized and submissive, whereas male-coded voices are associated with dominance and positive attributes, thereby reproducing existing gendered power structures. By integrating social critique into the evaluation of speech synthesis technologies, this work advances a more inclusive and socially aware approach to voice AI design.
📝 Abstract
This work examines sexualised AI-generated English-speaking voices offered by a popular commercial platform. New technologies may enable sexual empowerment and greater diversity in gender expression, yet toxic masculinity, heteronormativity, and the abuse of women and LGBTQ+ people remain pervasive online. Drawing on a Feminist HCI perspective, we examine how commercial voice AI systems reproduce and circulate particular performances of gender. We conducted a listening experiment with a diverse group of listeners, combining quantitative adjective selection, qualitative free-text responses, and acoustic analysis. Participants evaluated male- and female-coded voices presented with either sexualised scripts or neutral text. Results reveal a narrow range of gender expression, largely binary and heteronormative. Female-coded voices are more frequently described using sexualised and submissive terms, while male-coded voices are more often associated with dominance and positive traits.