ArGuard Shared Task: Harmful Content Detection in Arabic Memes and LLM Prompts

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of detecting harmful content in Arabic memes and large language model (LLM) prompts by organizing a shared task comprising 35 teams across two tracks: multimodal hate speech detection and LLM safety evaluation. Methodologically, it introduces the first fine-grained benchmark for Arabic harmful content to expose distribution shift challenges, and systematically evaluates state-of-the-art models including AraBERT, Jais, and Qwen3-VL. The top-performing systems achieved macro-F1 scores ranging from 0.823 to 0.984, establishing robust performance baselines for this domain. Ultimately, this work provides critical data resources and methodological references to advance AI safety research in the Arabic language context.
📝 Abstract
ArGuard is a shared task on harmful content detection in Arabic memes and LLM prompts. It includes two tracks: Track A focuses on multimodal hate detection in Arabic memes, while Track B addresses harmful prompt detection for Arabic LLM safety evaluation. In total, 58 teams registered, 35 participated in the final evaluation, and 27 submitted system-description papers. Participating teams explored models such as AraBERT, Jais, and Qwen3-VL. The best systems achieved macro-F1 scores of 0.823 on A1, 0.419 on A2, 0.984 on B1, and 0.790 on B2. Fine-grained meme classification in A2 was the most challenging setting, partly due to sparse labels and train-test distribution shifts.
Problem

Research questions and friction points this paper is trying to address.

Harmful Content Detection
Arabic Memes
LLM Safety
Multimodal Hate Detection
Harmful Prompt Detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Harmful Content Detection
Arabic Memes
Multimodal Hate Detection
LLM Safety
Prompt Detection
🔎 Similar Papers
No similar papers found.