Learn Before You Judge: Progressive Knowledge-to-Decision Alignment for Explainable Hateful Meme Detection

📅 2026-09-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决仇恨表情包检测中的任务干扰问题,提出ProKDA方法,通过逐步学习背景知识、检测学习及边界对齐来提高检测性能和解释性。
📝 Abstract
Hateful memes spread abusive content through implicit interactions between images and text, posing serious threats to the safety of online communities. In recent years, multimodal large language models have been widely used for hateful meme detection and are increasingly adopted to generate explainable detection results. However, we find that existing explain-then-detect methods often couple explanation generation and label prediction within the same training process. This coupling causes interference between task objectives, leading to limited detection performance and even worse results than simple SFT baselines. To address these challenges, we propose ProKDA, a progressive knowledge-to-decision alignment method for explainable hateful meme detection. Inspired by the human annotation training process, ProKDA first uses an agentic background knowledge construction pipeline to obtain external knowledge related to meme understanding. It then adopts a three-stage training strategy that sequentially performs background knowledge learning, hatefulness detection learning, and hatefulness boundary alignment. Unlike prior explain-then-detect methods that jointly optimize both tasks, ProKDA focuses on a single training objective at each stage. This design reduces interference between the two tasks and progressively transforms background knowledge into robust detection decisions. Experiments on three public hateful meme benchmarks show that ProKDA achieves state-of-the-art detection performance and provides accurate, explainable, and evidence-supported decisions for hateful meme moderation. Project page: https://meizhiyuan88666.github.io/prokda.
Problem

Research questions and friction points this paper is trying to address.

hateful meme detection
explain-then-detect methods
task interference
detection performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Progressive Knowledge-to-Decision Alignment
Explainable Hateful Meme Detection
Three-Stage Training Strategy
Bo Xu
Bo Xu
Dalian University of Technology
Natural Language ProcessingInformation RetrievalMedical DialoguePsychological Computing
C
Chenyuan Wang
School of Software, Dalian University of Technology
X
Xinyu Chen
School of Software, Dalian University of Technology
Q
Quanhao Zhu
School of Software, Dalian University of Technology
R
Rui Lin
School of Software, Dalian University of Technology
L
Liang Zhao
School of Software, Dalian University of Technology
Hongfei Lin
Hongfei Lin
DalianUniversity of Technology
natural language processing,sentimental analysistext miningsocial computing
Feng Xia
Feng Xia
Professor, School of Computing Technologies, RMIT University
Artificial IntelligenceGraph LearningBrainRoboticsCyber-Physical Systems