Automatic Evaluation of Mental Health Stigma in Online Communication

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of automatically evaluating mental health stigma given its complex manifestations. To this end, we construct a theory-driven, multi-level taxonomy encompassing stigma modes, domains, and components, and perform fine-grained annotation on news and social media texts to establish a multi-disorder stigma detection benchmark. Methodologically, we employ large language model (LLM) evaluation techniques alongside sentiment, toxicity, and hate speech classifiers to conduct comparative experiments and rule optimization. Our findings demonstrate that general-purpose models fail to effectively identify mental health stigma, revealing critical limitations in existing approaches. Furthermore, this work highlights the pivotal role of explicit operational rules in enhancing LLM predictive accuracy for stigma detection tasks.
📝 Abstract
Mental health stigma has profoundly harmful impacts but its complexity makes it difficult to evaluate. Stigma may involve explicit derogation, but also subtler forms of blame, fear, paternalistic pity, social distancing, structural exclusion, and discrimination. We introduce a theory-grounded benchmark for automatic evaluation of mental health stigma in online communication, consisting of naturally occurring online news and social media text annotated with a fine-grained taxonomy of stigma across multiple mental health conditions. Our annotation framework comprises a binary stigma-detection task and a multi-level taxonomy covering (i) stigma mode, (ii) domain, and (iii) specific components of certain forms of stigma. We apply this framework to texts mentioning six mental health conditions and evaluate large language models alongside stigma-related classifiers for detecting sentiment, toxicity, and hate speech. Results show that mental health stigma is not well captured by models trained to detect these neighboring constructs, and that LLMs often overpredict stigma unless given explicit operational rules - mirroring the importance of decision rules in human annotation. We release the publicly available part of benchmark, annotations, prototypical exemplars of stigma and code at: https://github.com/jemimakang/mh_stigma.
Problem

Research questions and friction points this paper is trying to address.

mental health stigma
automatic evaluation
online communication
large language models
benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mental Health Stigma
Benchmark
Large Language Models
Fine-grained Taxonomy
Automatic Evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
N
Naomi Baes
Melbourne School of Psychological Sciences, The University of Melbourne; Department of General and Computational Linguistics, University of Tübingen
J
Jemima Kang
School of Computing and Information Systems, The University of Melbourne
Nick Haslam
Nick Haslam
Melbourne School of Psychological Sciences, The University of Melbourne
C
Chris Groot
Melbourne School of Psychological Sciences, The University of Melbourne
A
Alsa Wu
Melbourne School of Psychological Sciences, The University of Melbourne
L
Luc Raszewski
School of Computing and Information Systems, The University of Melbourne
Yulia Otmakhova
Yulia Otmakhova
Research Fellow grade 2 (eq. Assistant Professor), University of Melbourne
NLPbioNLP