🤖 AI Summary
This study systematically evaluates ChatGPT (particularly version 6) in detecting inappropriate and targeted language within social media user-generated content (UGC). We employ zero-shot and few-shot prompting strategies and construct a multi-source, human-annotated benchmark—combining crowdsourced and expert annotations—augmented by cross-level consistency analysis and error attribution to quantify model accuracy, coverage, and stability. Results show a significant improvement in inappropriate language detection accuracy; however, targeted language detection achieves only an F1-score of 0.72, with a false positive rate 18 percentage points higher than expert annotators—revealing critical limitations in contextual and intent understanding. To our knowledge, this is the first work to empirically characterize the performance divergence between these two closely related content moderation tasks. We further propose context-enhanced prompting and iterative fine-tuning as viable optimization pathways. The study delivers a reproducible evaluation framework and empirically grounded operational boundaries for AI-assisted content moderation.
📝 Abstract
This study evaluates the effectiveness of ChatGPT, an advanced AI model for natural language processing, in identifying targeting and inappropriate language in online comments. With the increasing challenge of moderating vast volumes of user-generated content on social network sites, the role of AI in content moderation has gained prominence. We compared ChatGPT's performance against crowd-sourced annotations and expert evaluations to assess its accuracy, scope of detection, and consistency. Our findings highlight that ChatGPT performs well in detecting inappropriate content, showing notable improvements in accuracy through iterative refinements, particularly in Version 6. However, its performance in targeting language detection showed variability, with higher false positive rates compared to expert judgments. This study contributes to the field by demonstrating the potential of AI models like ChatGPT to enhance automated content moderation systems while also identifying areas for further improvement. The results underscore the importance of continuous model refinement and contextual understanding to better support automated moderation and mitigate harmful online behavior.