Quantifying and Predicting Disagreement in Graded Human Ratings

📅 2026-05-01
📈 Citations: 0
Influential: 0
📄 PDF

career value

162K/year
🤖 AI Summary
This work addresses the substantial and sample-dependent disagreement among human annotators when labeling inappropriate language, such as offensive or hateful content—a phenomenon that is difficult to quantify. The authors propose an “opposition index” to characterize the degree of annotator polarization and develop methods to predict this index and the associated annotation variance based on textual features. They systematically compare two approaches: direct regression to predict variance and variance estimation derived from predicted probability distributions. Both achieve moderate predictive performance. The study further reveals that samples with high opposition indices are more challenging for models to classify accurately and tend to have their toxicity systematically underestimated. This research offers a novel perspective and practical tools for understanding and modeling subjective annotation disagreement in toxic language detection.
📝 Abstract
It is increasingly recognized that human annotators do not always agree, and such disagreement is inherent in many annotation tasks. However, not all instances in a given task elicit the same degree of opinion divergence. In this paper, we investigate annotation variation patterns in graded human ratings for inappropriate languages, including offensive language, hate speech, and toxic language perception. We examine whether the degree of annotation disagreement can be predicted from textual features. We further propose the Opposition Index, a metric that quantifies perspective opposition among annotators on a given item, and investigate the predictability of instances with potentially opposing human opinions. Our results show a moderate positive correlation between estimated and observed annotation variance. We find that two approaches achieve comparable performance in variance prediction: directly predicting the variance value and estimating it from predicted annotation distributions. Our results on opposition perspective prediction show that items with high opposition index values are more difficult to predict and are often underestimated by models.
Problem

Research questions and friction points this paper is trying to address.

annotation disagreement
graded ratings
offensive language
hate speech
toxic language
Innovation

Methods, ideas, or system contributions that make the work stand out.

Opposition Index
annotation disagreement
graded human ratings
variance prediction
toxic language perception
🔎 Similar Papers
No similar papers found.