Towards participatory speech dataset curation: A queer case study and conceptual framework

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文通过LGBTQIA+社区案例,提出了一种参与式语音数据集创建框架,以解决AI对边缘化群体的潜在危害问题。
📝 Abstract
In this paper, we motivate the need for a participatory speech dataset creation framework through a case study of the LGBTQIA+, or queer, community - a community with documented concerns about AI and reported harms, including attempts to develop 'gaydar' technologies that purportedly identify individuals as queer. We review common speech data collection practices, why these methods may be unsuitable for engaging with queer speakers, and discuss previous efforts in participatory AI with queer community engagement, as well as participatory endeavours specific to speech data collection for other marginalized communities. From this review, we develop a conceptual framework for participatory speech data curation by, for, and with marginalized communities drawing on insights from co-design and knowledge sharing. We propose a framework comprising overlapping and two-way processes of defining a community, project formulation, modes of participation, and personal autonomy.
Problem

Research questions and friction points this paper is trying to address.

participatory speech dataset
queer community
marginalized communities
Innovation

Methods, ideas, or system contributions that make the work stand out.

participatory speech dataset
queer community
co-design
knowledge sharing
marginalized communities
🔎 Similar Papers
No similar papers found.