Hatebench in the era of safer LLMs

📅 2026-09-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文研究了现有仇恨言论检测器对大语言模型生成的仇恨内容的泛化能力,并评估了其随时间和技术进步的稳定性,发现新模型已采取措施防止生成有害内容。
📝 Abstract
As Large Language Models (LLMs) lower the barrier for au- tomated content generation, the potential for producing hate speech poses a significant challenge for digital safety. This paper presents a reproducibility study of the HateBench paper by Shen et al., investigating whether existing hate speech detectors, typically trained on human-authored data, generalize to LLM-generated hateful content, and evaluating whether their reported weaknesses are stable over time and robust to evolving components. We independently reconstruct the original dataset genera- tion pipeline using modern LLMs and extend the benchmark to include recently released models and updated detector versions. Our independent assessment under current con- ditions finds that for newer LLMs, safeguards have been put into place to prevent the generation of harmful content. We also replicate the results for two sophisticated types of hate campaigns. While the original findings seem to have been overestimated slightly due to bias in the datasets, the overall findings can be confirmed. Finally, we compare text- Moderation against the newer omni-Moderation and find that its robustness against adversarial hate campaigns has improved slightly. By clarifying which detector vulnerabil- ities persist, this study informs the community about the longevity of content moderation measurements.
Problem

Research questions and friction points this paper is trying to address.

Hate speech
Large Language Models
Content Moderation
Digital Safety
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
hate speech detectors
omni-Moderation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
O
Ole Becker
Hasso Plattner Institute, University of Potsdam
T
Tobias Jongen
Hasso Plattner Institute, University of Potsdam
P
Philip Kolbe
Hasso Plattner Institute, University of Potsdam
S
Sonal Khosla
Hasso Plattner Institute, University of Potsdam
V
Vaibhav Bajpai
Hasso Plattner Institute, University of Potsdam