Large Language Models' Varying Accuracy in Recognizing Risk-Promoting and Health-Supporting Sentiments in Public Health Discourse: The Cases of HPV Vaccination and Heated Tobacco Products

📅 2025-07-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Large language models (LLMs) exhibit uncharacterized biases in public health sentiment analysis, particularly for nuanced emotional categories such as risk-promoting and health-supportive sentiments. Method: This study conducts a systematic, cross-platform (Facebook vs. Twitter), cross-topic (HPV vaccination vs. heated tobacco products), and cross-model (GPT, Gemini, LLaMA) evaluation, using human-annotated real-world social media data as the gold standard. Contribution/Results: All three LLMs achieve high overall accuracy but demonstrate significant contextual dependency: Facebook yields superior performance for risk-promoting sentiment detection, whereas Twitter favors health-supportive sentiment identification; neutral sentiment remains consistently under-detected across platforms and models. Critically, this work is the first to empirically identify and characterize a three-dimensional bias pattern—spanning platform, health topic, and model architecture—in LLM-based public health sentiment analysis. These findings provide empirical guidance for LLM selection, domain adaptation, and bias mitigation in public health informatics.

Technology Category

Natural Language Processing: Ethics — Bias, Fairness, Transparency & PrivacyMachine Learning: Large Multimodal Models (LMMs)Computer Vision: Bias, Fairness & Privacy

Application Category

Social Networks and Social Media: Fairness and bias in social network and social media analysisUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Machine learning methods are increasingly applied to analyze health-related public discourse based on large-scale data, but questions remain regarding their ability to accurately detect different types of health sentiments. Especially, Large Language Models (LLMs) have gained attention as a powerful technology, yet their accuracy and feasibility in capturing different opinions and perspectives on health issues are largely unexplored. Thus, this research examines how accurate the three prominent LLMs (GPT, Gemini, and LLAMA) are in detecting risk-promoting versus health-supporting sentiments across two critical public health topics: Human Papillomavirus (HPV) vaccination and heated tobacco products (HTPs). Drawing on data from Facebook and Twitter, we curated multiple sets of messages supporting or opposing recommended health behaviors, supplemented with human annotations as the gold standard for sentiment classification. The findings indicate that all three LLMs generally demonstrate substantial accuracy in classifying risk-promoting and health-supporting sentiments, although notable discrepancies emerge by platform, health issue, and model type. Specifically, models often show higher accuracy for risk-promoting sentiment on Facebook, whereas health-supporting messages on Twitter are more accurately detected. An additional analysis also shows the challenges LLMs face in reliably detecting neutral messages. These results highlight the importance of carefully selecting and validating language models for public health analyses, particularly given potential biases in training data that may lead LLMs to overestimate or underestimate the prevalence of certain perspectives.
Problem

Research questions and friction points this paper is trying to address.

Assessing LLM accuracy in detecting health-related sentiments
Comparing LLM performance on risk-promoting vs health-supporting content
Identifying platform-specific biases in LLM sentiment classification
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evaluates three LLMs for sentiment detection accuracy
Uses human-annotated social media data as benchmark
Identifies platform and health topic accuracy variations
🔎 Similar Papers
No similar papers found.