🤖 AI Summary
This study addresses the widespread conflation of hostile content in state-sponsored social media influence operations under the umbrella term “hate,” which leads to substantial overestimation. To remedy this, the authors introduce the concept of “manufactured divisiveness” and develop an auditable classification framework that distinguishes identity-based hate speech from partisan attacks and geopolitical trolling. Combining a large language model–based dual-prompt detector with an expert rule system—and validated through human annotation (Cohen’s κ = 0.82/0.52)—the framework enables fine-grained analysis of influence campaigns across seven countries. Results reveal that only 18.7% of flagged content constitutes genuine hate speech, with the majority reflecting strategic divisiveness. The seven campaigns cluster into three distinct strategic types, demonstrating that conventional “hate” metrics obscure critical differences and have previously inflated estimates by approximately twofold.
📝 Abstract
State-backed influence operations are routinely measured as high-prevalence sources of ``hate'' and ``toxicity.'' We argue those rates rest on a measurement error: the detectors behind them are validated to catch a broader definition inclusive of hostility or divisiveness aimed at an out-group, and so over-attribute hate to content better described as partisan or geopolitical invective. Across 25.08M tweets from seven government-attributed campaigns in the Twitter Information Operations archive (8,275 accounts), we separate hate from the other forms of divisiveness. We first validate a two-prompt LLM-based detector, matching human labels at Cohen's $κ=0.82$, to identify the broader hostility; we then develop an auditable rule, agreeing with an expert at $κ=0.52$, to further classify this content (5,457 posts) into three sub-categories. About 50.1% are identity-based attacks on people, whereas 30.4% are partisan attacks and 19.5% invective against states and their foreign policy. Reporting all of it as hate therefore overstates hate roughly twofold; only 18.7% is both identity-based and dehumanizing or inciting. Six of seven campaigns sort into three regimes that a single ``hate'' rate flattens, namely identity hate (RU-op and IRA, both Russia-attributed), geopolitical invective (both Iran operations), and partisan divisiveness (both Venezuela operations). We call the shared product $manufactured divisiveness$. The line to separate these constructs itself remains unsettled: on the hardest cases three independent human experts agree only moderately (pairwise $κ=0.37$--$0.50$), and the best of nineteen LLM models tops out at $κ=0.601$ against the experts' majority. Our findings can help redefine the study of hate in the context of influence campaigns and broader online discourse.