Assessing and Refining ChatGPT's Performance in Identifying Targeting and Inappropriate Language: A Comparative Study

📅 2025-05-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study systematically evaluates ChatGPT (particularly version 6) in detecting inappropriate and targeted language within social media user-generated content (UGC). We employ zero-shot and few-shot prompting strategies and construct a multi-source, human-annotated benchmark—combining crowdsourced and expert annotations—augmented by cross-level consistency analysis and error attribution to quantify model accuracy, coverage, and stability. Results show a significant improvement in inappropriate language detection accuracy; however, targeted language detection achieves only an F1-score of 0.72, with a false positive rate 18 percentage points higher than expert annotators—revealing critical limitations in contextual and intent understanding. To our knowledge, this is the first work to empirically characterize the performance divergence between these two closely related content moderation tasks. We further propose context-enhanced prompting and iterative fine-tuning as viable optimization pathways. The study delivers a reproducible evaluation framework and empirically grounded operational boundaries for AI-assisted content moderation.

Technology Category

Natural Language Processing: Fact-Checking / Misinformation Detection (NLP Focus)Machine Learning: Large Multimodal Models (LMMs)Computer Vision: Language and Vision

Application Category

Social Networks and Social Media: Generative AI / large language models and their impact on social systemsWeb Mining and Content Analysis: Large pretrained models with web dataSearch and Retrieval-Augmented AI: Web evaluation methodologies and metrics
📝 Abstract
This study evaluates the effectiveness of ChatGPT, an advanced AI model for natural language processing, in identifying targeting and inappropriate language in online comments. With the increasing challenge of moderating vast volumes of user-generated content on social network sites, the role of AI in content moderation has gained prominence. We compared ChatGPT's performance against crowd-sourced annotations and expert evaluations to assess its accuracy, scope of detection, and consistency. Our findings highlight that ChatGPT performs well in detecting inappropriate content, showing notable improvements in accuracy through iterative refinements, particularly in Version 6. However, its performance in targeting language detection showed variability, with higher false positive rates compared to expert judgments. This study contributes to the field by demonstrating the potential of AI models like ChatGPT to enhance automated content moderation systems while also identifying areas for further improvement. The results underscore the importance of continuous model refinement and contextual understanding to better support automated moderation and mitigate harmful online behavior.
Problem

Research questions and friction points this paper is trying to address.

Evaluating ChatGPT's accuracy in detecting inappropriate online language
Comparing AI performance with human annotations for content moderation
Identifying variability in targeting language detection by ChatGPT
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evaluates ChatGPT for inappropriate language detection
Compares performance with crowd and expert annotations
Highlights iterative refinements improve accuracy
🔎 Similar Papers
No similar papers found.
B
Barbarestani Baran
Vrije Universiteit Amsterdam, Computational Linguistics and Text Mining Lab
M
Maks Isa
Vrije Universiteit Amsterdam, Computational Linguistics and Text Mining Lab
V
Vossen Piek
Vrije Universiteit Amsterdam, Computational Linguistics and Text Mining Lab