Nobody Truly Agrees on Sentiment: Humans, Bespoke Tools, and LLMs Struggle with Social Media Texts

πŸ“… 2026-10-07
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenges of insufficient evaluation reliability and inherent subjectivity in social media sentiment analysis tools by systematically comparing inter-rater agreement among human annotators, specialized sentiment analyzers, and large language models (LLMs) for tweet sentiment classification. Employing Cohen’s Kappa and Fleiss’ Kappa statistics, the research conducts a quantitative multi-source evaluation. Results indicate that agreement is significantly higher in binary classification than in fine-grained tasks. Notably, the domain-adapted Twitter-RoBERTa model most closely approximates human judgment, whereas general-purpose LLMs perform inadequately. This work highlights the subjectivity bottleneck in sentiment annotation, validates the critical role of domain-specific fine-tuning, and underscores the necessity of incorporating human evaluation to ensure the reliability of sentiment analysis systems.
πŸ“ Abstract
Social media is a rich source of real-time public sentiment, but widely used sentiment analysis tools are often applied without understanding their limitations. In this study, we evaluate the inter-rater reliability of three bespoke sentiment analysis tools (TextBlob, VADER, and Twitter-roBERTa-base) and three large language models (LLMs: Qwen3-32B, GPT-OSS-120B, Llama-4-Maverick-17B) against six human raters across 100 tweets. We measured agreement using two statistical measures: Cohen's kappa for pairwise comparisons and Fleiss' kappa for multiple raters. Even among the human raters, our results showed only fair agreement, highlighting the subjectivity of sentiment analysis. Higher agreement was observed under the binary sentiment classification (negative vs. non-negative and positive vs. non-positive) than under the three-class classification across both humans and automated tools. The Twitter-roBERTa-base model showed the strongest alignment with human ratings, outperforming both bespoke sentiment tools and LLMs, particularly in distinguishing negative versus non-negative sentiment. LLMs showed substantial agreement among themselves and moderate to substantial alignment with humans, performing better in positive vs. non-positive classifications. Our findings underscore that domain-specific fine-tuning remains crucial for reliable social media sentiment analysis, and human-centered evaluation remains essential for establishing gold-standard labels.
Problem

Research questions and friction points this paper is trying to address.

Sentiment Analysis
Social Media
Inter-rater Reliability
Large Language Models
Subjectivity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sentiment Analysis
Inter-rater Reliability
Large Language Models
Domain-specific Fine-tuning
Social Media
πŸ”Ž Similar Papers
No similar papers found.