π€ AI Summary
This study addresses the challenges posed by the proliferation of online content and the exacerbation of hate speech generation by large language models (LLMs), highlighting the urgent need for efficient, adaptable classification schemes in existing moderation systems. We propose HATEDECIDE, an evaluation framework that systematically compares six structured decision-model configurations against multiple baselines to investigate whether providing explicit definitions or decomposing tasks yields practical gains for zero-shot hate speech detection. Experimental results demonstrate that the optimal hosted model approximates the performance of commercial LLMs while reducing inference costs by approximately 97%; however, explicit criteria do not necessarily improve classification accuracy. This work provides empirical evidence supporting low-cost, configurable automated content moderation.
π Abstract
The scale of online content makes hate-speech moderation challenging, while Large Language Models (LLMs) enable harmful material to be produced and adapted more easily. Moderation therefore requires efficient classifiers that can accommodate different definitions of hate speech. Recent structured decision models accept natural-language criteria and select among specified answers, raising the question of whether they can meet these requirements without task-specific training. We present HATEDECIDE, an evaluation of six decision-model configurations on four hate-speech datasets against specialized moderation, zero-shot, commercial, and supervised baselines. We examine whether supplying a dataset's definition, or decomposing it into multiple questions, improves classification, and we measure their latency and cost. We find that commercial LLMs significantly outperform all decision models on only one dataset. Supplying definitions changes up to 28\% of predictions without consistently improving classification, and decomposition significantly improves performance in only 20\% of the comparisons. On a diagnostic set of test cases, the best hosted decision model comes within 1.6 macro-F1 points of the best commercial LLM at approximately 97\% lower inference cost. These results identify opportunities for inexpensive moderation, while showing that explicit criteria and additional questions do not reliably improve classification.