ArGuard Shared Task: Harmful Content Detection in Arabic Memes and LLM Prompts
This study addresses the challenges of detecting harmful content in Arabic memes and large language model (LLM) prompts by organizing a shared task comprising 35 teams across two tracks: multimodal hate speech detection and LLM safety evaluation. Methodologically, it introduces the first fine-grained benchmark for Arabic harmful content to expose distribution shift challenges, and systematically evaluates state-of-the-art models including AraBERT, Jais, and Qwen3-VL. The top-performing systems achieved macro-F1 scores ranging from 0.823 to 0.984, establishing robust performance baselines for this domain. Ultimately, this work provides critical data resources and methodological references to advance AI safety research in the Arabic language context.