Large Language Models' Accuracy in Emulating Human Experts' Evaluation of Public Sentiments about Heated Tobacco Products on Social Media

📅 2025-01-31
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study systematically evaluates the capability of large language models (LLMs) to replicate human experts’ fine-grained sentiment judgments—anti-HTP, neutral, or pro-HTP—in social media discourse on heated tobacco products (HTPs). Using extensive prompt engineering and majority-voting ensemble methods based on GPT-3.5 and GPT-4 Turbo, we conduct consistency assessments against human-annotated gold-standard labels on Facebook and Twitter data. Our results demonstrate, for the first time, that GPT-4 Turbo achieves 99% of the accuracy attained by 20-sample sampling with only three samples; it attains significantly higher accuracy (81.7%/77.0% on Facebook/Twitter) than GPT-3.5 (61.2%/57.0%) and exhibits greater robustness to biased or polarized text. This work establishes a reproducible, high-accuracy methodological paradigm and empirical benchmark for LLM-driven automated sentiment monitoring in public health applications.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsHumans and AI: Human-in-the-loop Machine Learning

Application Category

Economics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Large language models for search
📝 Abstract
Sentiment analysis of alternative tobacco products on social media is important for tobacco control research. Large Language Models (LLMs) can help streamline the labor-intensive human sentiment analysis process. This study examined the accuracy of LLMs in replicating human sentiment evaluation of social media messages about heated tobacco products (HTPs). The research used GPT-3.5 and GPT-4 Turbo to classify 500 Facebook and 500 Twitter messages, including anti-HTPs, pro-HTPs, and neutral messages. The models evaluated each message up to 20 times, and their majority label was compared to human evaluators. Results showed that GPT-3.5 accurately replicated human sentiment 61.2% of the time for Facebook messages and 57.0% for Twitter messages. GPT-4 Turbo performed better, with 81.7% accuracy for Facebook and 77.0% for Twitter. Using three response instances, GPT-4 Turbo achieved 99% of the accuracy of twenty instances. GPT-4 Turbo also had higher accuracy for anti- and pro-HTPs messages compared to neutral ones. Misclassifications by GPT-3.5 often involved anti- or pro-HTPs messages being labeled as neutral or irrelevant, while GPT-4 Turbo showed improvements across all categories. In conclusion, LLMs can be used for sentiment analysis of HTP-related social media messages, with GPT-4 Turbo reaching around 80% accuracy compared to human experts. However, there's a risk of misrepresenting overall sentiment due to differences in accuracy across sentiment categories.
Problem

Research questions and friction points this paper is trying to address.

Assessing LLMs' accuracy in sentiment analysis.
Comparing GPT-3.5 and GPT-4 Turbo performance.
Evaluating sentiment on heated tobacco products.
Innovation

Methods, ideas, or system contributions that make the work stand out.

GPT-4 Turbo sentiment analysis
Replicates human evaluation accuracy
Classifies social media messages effectively
🔎 Similar Papers
No similar papers found.