🤖 AI Summary
This study addresses the challenge of hate speech detection in Mexican Spanish videos arising from the scarcity of non-English resources. To this end, it introduces MexHat, the first multimodal dataset tailored to this linguistic context. The dataset comprises one thousand meticulously human-annotated videos encompassing a three-level classification taxonomy and fine-grained categories. By integrating multimodal content analysis techniques, the proposed approach effectively captures language-specific cultural cues. Beyond establishing a comprehensive benchmark, this work elucidates the inherent challenges of hate speech detection across diverse cultural contexts. Ultimately, the introduction of MexHat significantly enhances the cultural and contextual awareness of multimodal models, facilitating more robust cross-cultural hate speech recognition.
📝 Abstract
Ensuring online safety through content monitoring had raised Hate Speech Detection as a crucial task to be addressed. By essence the task demands the capture of contextual cues, which are essential for a precise understanding of the content's intent. Although automated detection approaches for the task have advanced significantly, the scarcity of non-English resources persists, limiting the ability of models to adapt to the subtle, context-dependent, and culturally related nature of multimodal content. In this paper, we introduce MexHat, a video dataset designed to capture the linguistic and cultural cues for the hate-speech detection task in a Mexican Spanish context. Our dataset comprises around 1k video clips annotated across two tasks: a three-way class evaluation (no negative content, offensive content and hate-speech content), and a fine-grained class evaluation including three hate-speech sub-categories. The dataset statistics and the baseline results highlight the inherent challenges associated with the task. Disclaimer: This paper contains sensitive content that may be disturbing to some readers.