What is Beneath Misogyny: Misogynous Memes Classification and Explanation

πŸ“… 2025-07-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the detection of implicit, multimodal (image-text co-dependent), and socially contextualized misogyny in internet memes. We propose MM-Misogyny, a novel framework that first performs fine-grained cross-modal fusion via independent image and text encoders augmented with cross-attention mechanisms; it then leverages large language models to generate interpretable classification rationales and behavioral attributions. To support this work, we introduce WBMSβ€”the first multicontextual, expert-annotated benchmark dataset specifically designed for misogynistic memes. Extensive experiments demonstrate that MM-Misogyny significantly outperforms existing baselines on both binary detection and fine-grained classification tasks. The code and dataset are publicly released, establishing a new paradigm for explainable, multimodal analysis of digital gender bias.

Technology Category

Machine Learning: Multimodal LearningComputer Vision: Multi-modal VisionNatural Language Processing: Fact-Checking / Misinformation Detection (NLP Focus)

Application Category

Web Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataSocial Networks and Social Media: Fairness and bias in social network and social media analysisSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web search
πŸ“ Abstract
Memes are popular in the modern world and are distributed primarily for entertainment. However, harmful ideologies such as misogyny can be propagated through innocent-looking memes. The detection and understanding of why a meme is misogynous is a research challenge due to its multimodal nature (image and text) and its nuanced manifestations across different societal contexts. We introduce a novel multimodal approach, extit{namely}, extit{ extbf{MM-Misogyny}} to detect, categorize, and explain misogynistic content in memes. extit{ extbf{MM-Misogyny}} processes text and image modalities separately and unifies them into a multimodal context through a cross-attention mechanism. The resulting multimodal context is then easily processed for labeling, categorization, and explanation via a classifier and Large Language Model (LLM). The evaluation of the proposed model is performed on a newly curated dataset ( extit{ extbf{W}hat's extbf{B}eneath extbf{M}isogynous extbf{S}tereotyping (WBMS)}) created by collecting misogynous memes from cyberspace and categorizing them into four categories, extit{namely}, Kitchen, Leadership, Working, and Shopping. The model not only detects and classifies misogyny, but also provides a granular understanding of how misogyny operates in domains of life. The results demonstrate the superiority of our approach compared to existing methods. The code and dataset are available at href{https://github.com/kushalkanwarNS/WhatisBeneathMisogyny/tree/main}{https://github.com/Misogyny}.
Problem

Research questions and friction points this paper is trying to address.

Detect and classify misogynous content in memes
Explain nuanced misogyny across societal contexts
Analyze multimodal meme data (image and text)
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal approach for meme analysis
Cross-attention mechanism for text-image fusion
LLM-based explanation for misogyny detection