Beyond the pale: Assessing prevalence and contents of extremist speech in LLM training data

📅 2026-08-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the critical gap in assessing extremist content within large language model pretraining corpora by systematically quantifying safety risks in the open-source Dolma dataset. We propose an automated detection framework that integrates multi-source definitions with expert validation to overcome the limitations of single-standard approaches. Empirical analysis reveals hundreds of thousands of instances of hate speech and violent content, confirming significant safety hazards inherent in current pretraining data. Beyond characterizing the prevalence and composition of extremist rhetoric, this work provides essential empirical evidence and methodological foundations for advancing data governance and safety alignment in large language models.
📝 Abstract
Despite a strong interest on the part of the research community in the topic of trustworthy and safe AI, the composition of the text corpora that large language models (LLMs) encounter in pre- and post-training has not yet drawn much attention. In this work, we address the question of whether LLMs are exposed to unfiltered, uncontextualised extremist speech. Using several definitions of extremist speech, stemming from official documents and research literature, and an extraction pipeline combining automated text processing with expert verification, we provide a lower bound on the prevalence of extremist documents in Dolma, an open training corpus underpinning the OLMo series of models. We show that Dolma is likely to include hundreds of thousands of documents containing extremist content and hate speech of several types, including direct calls for violence, and discuss the implications of this for data curation and model pre-training.
Problem

Research questions and friction points this paper is trying to address.

Extremist speech
LLM training data
Data curation
Hate speech
AI safety
Innovation

Methods, ideas, or system contributions that make the work stand out.

Extremist Speech Detection
Training Data Curation
Expert-in-the-loop Verification
LLM Safety Assessment
Dolma Corpus Analysis
🔎 Similar Papers
No similar papers found.