🤖 AI Summary
This study addresses the societal and operational harms arising from rapidly deployed AI systems, which remain inadequately captured by conventional pre-deployment threat modeling and post-hoc tracking. We propose a large language model-based thematic analysis pipeline that dynamically detects, categorizes, and longitudinally tracks harms through Reddit post summaries, constructing a bottom-up AI harm taxonomy. The project releases a dataset comprising 575,000 posts alongside a hierarchical taxonomy encompassing 12 primary categories and 47 sub-nodes. While aligning with expert-defined risk frameworks, this approach reveals nuanced harms overlooked by top-down schemas, such as agent-mediated privacy leakage. By substantially shortening the harm detection cycle, this methodology establishes a novel paradigm for participatory AI governance.
📝 Abstract
The rapid deployment of AI systems has created socio-technical, psychological, and operational harms that can elude ex-ante threat modelling and ex-post incident tracking. We introduce an LLM-assisted thematic analysis pipeline to dynamically detect, categorise, and track emerging AI harms from large-scale social media data. Applying it to 5.7 million Reddit post summaries over 18 months (01/2025 to 06/2026), we curate and release a dataset of 575,000 AI harm-related posts and a bottom-up AI harm taxonomy of 12 categories and 47 subnodes. The taxonomy reliably covers established expert-defined risks while surfacing granular harms that top-down frameworks overlook, such as distinct forms of AI privacy violations. Temporal analysis surfaces evolving user-centric harms, such as agentic privacy and security breaches, premature AI adoption in the workplace, and grief from AI companion discontinuation. Our pipeline shortens harm-detection timelines and hereby complements efforts towards more participatory and responsive AI governance.