π€ AI Summary
This study addresses the detection of implicit, multimodal (image-text co-dependent), and socially contextualized misogyny in internet memes. We propose MM-Misogyny, a novel framework that first performs fine-grained cross-modal fusion via independent image and text encoders augmented with cross-attention mechanisms; it then leverages large language models to generate interpretable classification rationales and behavioral attributions. To support this work, we introduce WBMSβthe first multicontextual, expert-annotated benchmark dataset specifically designed for misogynistic memes. Extensive experiments demonstrate that MM-Misogyny significantly outperforms existing baselines on both binary detection and fine-grained classification tasks. The code and dataset are publicly released, establishing a new paradigm for explainable, multimodal analysis of digital gender bias.
π Abstract
Memes are popular in the modern world and are distributed primarily for entertainment. However, harmful ideologies such as misogyny can be propagated through innocent-looking memes. The detection and understanding of why a meme is misogynous is a research challenge due to its multimodal nature (image and text) and its nuanced manifestations across different societal contexts. We introduce a novel multimodal approach, extit{namely}, extit{ extbf{MM-Misogyny}} to detect, categorize, and explain misogynistic content in memes. extit{ extbf{MM-Misogyny}} processes text and image modalities separately and unifies them into a multimodal context through a cross-attention mechanism. The resulting multimodal context is then easily processed for labeling, categorization, and explanation via a classifier and Large Language Model (LLM). The evaluation of the proposed model is performed on a newly curated dataset ( extit{ extbf{W}hat's extbf{B}eneath extbf{M}isogynous extbf{S}tereotyping (WBMS)}) created by collecting misogynous memes from cyberspace and categorizing them into four categories, extit{namely}, Kitchen, Leadership, Working, and Shopping. The model not only detects and classifies misogyny, but also provides a granular understanding of how misogyny operates in domains of life. The results demonstrate the superiority of our approach compared to existing methods. The code and dataset are available at href{https://github.com/kushalkanwarNS/WhatisBeneathMisogyny/tree/main}{https://github.com/Misogyny}.