๐ค AI Summary
This study addresses the limitations of existing misogynistic speech detection methods, which lack systematic annotation criteria grounded in psychological and philosophical theories, thereby failing to comprehensively capture online misogyny. To bridge this gap, the authors propose a novel annotation framework that integrates empirically validated psychological theories, resulting in a high-quality annotation guideline and dataset. The frameworkโs reliability is demonstrated through human annotation (Cohenโs kappa = 0.68) and evaluation via large language model (LLM) classification experiments. Results show that the proposed approach outperforms current expert-based methods across three datasets. Furthermore, the study reveals that LLMs struggle to replicate human judgments on misogynistic content due to their reliance on dominant societal narratives rather than theoretically informed frameworks, highlighting a critical limitation in their capacity for nuanced social harm detection.
๐ Abstract
Detecting misogynistic hate speech is a difficult algorithmic task. The task is made more difficult when decision criteria for what constitutes misogynistic speech are ungrounded in established literatures in psychology and philosophy, both of which have described in great detail the forms explicit and subtle misogynistic attitudes can take. In particular, the literature on algorithmic detection of misogynistic speech often rely on guidelines that are insufficiently robust or inappropriately justified -- they often fail to include various misogynistic phenomena or misrepresent their importance when they do. As a result, current misogyny detection coding schemes and datasets fail to capture the ways women experience misogyny online. This is of pressing importance: misogyny is on the rise both online and offline. Thus, the scientific community needs to have a systematic, theory informed coding scheme of misogyny detection and a corresponding dataset to train and test models of misogyny detection. To this end, we developed (1) a misogyny annotation guideline scheme informed by theoretical and empirical psychological research, (2) annotated a new dataset achieving substantial inter-rater agreement (kappa = 0.68) and (3) present a case study using Large Language Models (LLMs) to compare our coding scheme to a self-described"expert"misogyny annotation scheme in the literature. Our findings indicate that our guideline scheme surpasses the other coding scheme in the classification of misogynistic texts across 3 datasets. Additionally, we find that LLMs struggle to replicate our human annotator labels, attributable in large part to how LLMs reflect mainstream views of misogyny. We discuss implications for the use of LLMs for the purposes of misogyny detection.