Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks

📅 2026-05-11
📈 Citations: 0
Influential: 0
📄 PDF

career value

203K/year
🤖 AI Summary
Current benchmarks for evaluating toxicity in large language models exhibit underappreciated systematic biases that may lead to the deployment of unsafe models. This work systematically investigates how variations in task formulation—such as text completion versus summarization—input data domains, and evaluated models interact with multiple toxicity metrics. It reveals, for the first time, that both task type and data domain significantly influence toxicity scores. Experiments demonstrate that existing benchmarks are prone to misclassifying content as harmful when tasks are altered and show inconsistent performance across domains, highlighting their fragility and dependence on specific model-task configurations. These findings underscore the urgent need for more robust and reliable toxicity evaluation frameworks.
📝 Abstract
The rapid adoption of LLMs in both research and industry highlights the challenges of deploying them safely and reveals a gap in the systematic evaluation of toxicity benchmarks. As organizations increasingly rely on these benchmarks to certify models for customer-facing applications and automated moderation, unrecognized evaluation biases could lead to the deployment of vulnerable or unsafe systems. This work investigates the robustness of established benchmarking setups and examines how to measure currently neglected intrinsic biases, such as those related to model choice, metrics, and task types. Our experiments uncover significant discrepancies in benchmark behaviors when evaluation setups are altered. Specifically, shifting the task from text completion to summarization increases the tendency of benchmarks to flag content as harmful. Additionally, certain benchmarks fail to maintain consistent behavior when the input data domain is changed. Furthermore, we observe model-specific instabilities, demonstrating a clear need for more robust and comprehensive safety evaluation frameworks.
Problem

Research questions and friction points this paper is trying to address.

LLM evaluation
toxicity benchmarks
evaluation bias
safety evaluation
benchmark robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

toxicity benchmarks
evaluation bias
large language models
safety evaluation
intrinsic bias
🔎 Similar Papers
No similar papers found.
R
Regina Gugg
Dynatrace Research, Linz, Austria
S
Selina Niederländer
Dynatrace Research, Linz, Austria
A
Andreas Stöckl
University of Applied Sciences Upper Austria, Hagenberg, Austria
M
Martin Flechl
Dynatrace Research, Linz, Austria