Score
Designs, builds, and evaluates systems that identify, classify, and filter offensive language in text, including rule-based and machine-learning classifiers, context-aware models, and moderation pipelines. Manages and curates sensitive-word lists and policy rules, and analyzes performance tradeoffs (precision, recall, false positives/negatives) and operational behaviors for deployment and escalation.
The proliferation of online abusive language poses severe threats to individual and community well-being, necessitating a unified, actionable framework for detection and intervention. To address this, we propose the first hierarchical, multidimensional taxonomy of online abusive language, integrating annotation logics from 18 multilabel datasets. Our taxonomy systematically organizes 17 fine-grained dimensions across five core categories: context, target, intensity, directness, and theme. Methodologically, we combine systematic literature review, multilabel mapping, hierarchical clustering, and expert validation to ensure both theoretical rigor and practical scalability. The resulting open-source taxonomy has fostered initial consensus among researchers, platform operators, and policymakers on detection standards, cross-dataset alignment, and collaborative governance. It serves as a foundational tool for continuous monitoring, precise identification, and early intervention against online abuse.
Online textual abuse—including hate speech and cyberbullying—seriously harms users’ mental health and erodes social trust. While large language models (LLMs) enhance detection capabilities, they may also generate harmful content, exacerbating governance challenges. This study systematically reviews text abuse detection methods in Chinese social media and introduces, for the first time, a “technical–ethical” co-analysis framework. We empirically evaluate leading LLMs across four critical dimensions: detection accuracy, bias, robustness, and risk of generating abusive content. By integrating text classification, psychosocial impact modeling, and adversarial generation analysis, we uncover the dialectical role of LLMs—both mitigating and amplifying online abuse. Our findings provide empirically grounded, actionable insights for safe AI governance, including a phased technical roadmap for responsible deployment and mitigation.
This study addresses the urgent need for effective detection and neutralization of hate speech proliferating on social media. It systematically evaluates the performance of CNN, LSTM, BERT, and their variants in hate speech identification and proposes a novel text transformation method that automatically converts harmful content into semantically preserved neutral expressions. Furthermore, a hybrid model integrating the strengths of multiple architectures is developed, significantly enhancing detection accuracy in specific scenarios. Experimental results demonstrate that BERT-based models achieve superior performance owing to their deep contextual understanding, while the proposed text transformation strategy effectively mitigates the adverse impact of toxic content. The findings validate the feasibility and efficacy of a synergistic framework that jointly performs detection and neutralization.
This study systematically evaluates ChatGPT (particularly version 6) in detecting inappropriate and targeted language within social media user-generated content (UGC). We employ zero-shot and few-shot prompting strategies and construct a multi-source, human-annotated benchmark—combining crowdsourced and expert annotations—augmented by cross-level consistency analysis and error attribution to quantify model accuracy, coverage, and stability. Results show a significant improvement in inappropriate language detection accuracy; however, targeted language detection achieves only an F1-score of 0.72, with a false positive rate 18 percentage points higher than expert annotators—revealing critical limitations in contextual and intent understanding. To our knowledge, this is the first work to empirically characterize the performance divergence between these two closely related content moderation tasks. We further propose context-enhanced prompting and iterative fine-tuning as viable optimization pathways. The study delivers a reproducible evaluation framework and empirically grounded operational boundaries for AI-assisted content moderation.
The multi-type, overlapping nature of online hate speech renders conventional binary classification inadequate, motivating the shift toward multi-label classification. Method: We conduct the first systematic review of 46 English-language studies—spanning 28 datasets and 24 models—employing meta-analysis, cross-dataset consistency evaluation, and quantitative assessment of annotation quality (e.g., inter-annotator agreement, IAA). Contribution/Results: We reveal substantial heterogeneity in label taxonomies, dataset sizes, annotation rigor, and evaluation metrics. Key shared challenges include class imbalance, crowdsourcing bias, and sparse minority-label instances. Based on empirical findings, we propose 10 actionable methodological recommendations. We empirically validate the effectiveness of mainstream multi-label architectures—including BERT- and RNN-based models. This work establishes the first academic benchmark and practical guideline for developing robust, comparable, and regulation-compliant multi-label hate speech detection systems.
This work addresses offensive content detection in Hausa—a low-resource language—by constructing the first manually annotated Hausa offensive terminology dataset. Grounded in user surveys and empirical analysis, the study focuses on high-risk domains such as religion and politics to develop a domain-adapted detection system. Methodologically, it integrates supervised learning (XGBoost and fine-tuned multilingual BERT) with multilingual baselines, including direct translation via Google Translate. Key contributions are threefold: (1) release of the first open-source Hausa offensive dataset; (2) empirical demonstration that cultural context critically impacts detection performance, rendering literal translation ineffective; and (3) proposal of a localized, multi-stakeholder governance framework. Experiments show the proposed models achieve >70% accuracy—significantly outperforming translation-based baselines—and reveal pronounced concentration of offensive content in religious and political discourse.
Existing social media sensitive content detection tools suffer from limited customizability, narrow category coverage—particularly lacking long-tail classes such as drug-related and self-harm content—high privacy risks, and the absence of a unified evaluation benchmark. To address these issues, this work introduces the first high-quality, uniformly annotated dataset covering six sensitive content categories: conflict language, abuse, pornography, drug-related content, self-harm, and spam. We establish standardized protocols for data collection and human annotation. Leveraging this dataset, we supervise fine-tuning of open-source large language models (e.g., LLaMA) and design a comprehensive, multi-dimensional evaluation benchmark. Experimental results demonstrate that our approach consistently outperforms both the LLaMA baseline and the OpenAI API across all six detection tasks, achieving average improvements of 10–15%. Gains are especially pronounced for scarce categories (e.g., drug-related and self-harm content), validating the effectiveness and deployability of open-source LLM fine-tuning for fine-grained sensitive content identification.
This study investigates how political stance and cultural perspective influence large language models’ (LLMs) identification of offensive content in multilingual political tweets. Addressing the lack of ideological and cultural sensitivity in existing detection methods, we propose a personalized offensiveness assessment framework grounded in role-based prompting and chain-of-thought reasoning. We conduct systematic experiments across six mainstream LLMs—including DeepSeek-R1, Qwen3, and GPT-4.1-mini—on the MD-Agreement multilingual dataset. Results demonstrate that models with explicit reasoning capabilities achieve greater consistency and granularity in cross-lingual and cross-ideological settings; role-guided prompting significantly enhances modeling of cultural context and stance dependency. This work provides the first empirical evidence that reasoning mechanisms critically improve interpretability, judgment consistency, and personalized detection performance. It establishes a novel paradigm for value-aware NLP systems, advancing fairness and contextual fidelity in offensive language detection.
During sensitive periods—such as crises and elections—the detection of online harmful content (e.g., hate speech, offensive language) faces core challenges including conceptual ambiguity and poor generalizability across contexts and languages. Method: This study systematically reviews 140 relevant works to clarify definitional boundaries and data limitations; proposes a novel multilingual, cross-platform toxicity detection paradigm; and introduces a comprehensive benchmark dataset covering 32 languages and high-stakes scenarios—including elections and public health emergencies. Leveraging advanced machine learning and NLP techniques, the study optimizes classification models for enhanced cross-lingual and cross-platform robustness. Contribution/Results: The framework significantly improves accuracy and generalizability in toxic content identification, offering a reusable methodological foundation and empirically grounded guidelines for real-world content moderation practices.
This work addresses the limitations of existing abuse detection methods, which rely on static models and manual annotations and struggle to handle dynamic, context-sensitive online abusive behaviors. It proposes the first large language model (LLM)-integrated framework spanning the entire lifecycle of abuse detection, systematically encompassing label and feature generation, detection, appeal review, and audit governance. Bridging academic research and industrial practice, the study explores LLMs’ capabilities in contextual reasoning, policy interpretation, and cross-modal understanding. The authors delineate key architectural considerations for each phase, evaluate LLMs’ potential and limitations regarding interpretability, policy alignment, and multimodal fusion, and identify critical challenges—including latency, cost, determinism, adversarial robustness, and fairness—thereby offering a principled direction toward building reliable, accountable, large-scale abuse governance systems.
The proliferation of toxic text generated by large language models (LLMs) undermines the robustness of toxicity classifiers and increases their susceptibility to adversarial attacks. Method: This paper proposes a mechanistic interpretability–driven active defense framework: it introduces attention-head-level circuit analysis—the first such application—to diagnose classifier vulnerabilities; integrates fine-grained attribution with adversarial attack localization to identify critical, attack-prone components; and enhances robustness via targeted circuit suppression. Contribution/Results: Evaluated on BERT and RoBERTa architectures across diverse demographic datasets, the method significantly improves classification accuracy under adversarial perturbations. It further uncovers systematic differences in model vulnerability across demographic groups, revealing fairness-related failure modes. By unifying interpretability, robustness, and fairness, this work establishes a novel paradigm for building trustworthy, auditable, and attack-resilient content moderation systems.
This study addresses the challenges posed by the proliferation of online content and the exacerbation of hate speech generation by large language models (LLMs), highlighting the urgent need for efficient, adaptable classification schemes in existing moderation systems. We propose HATEDECIDE, an evaluation framework that systematically compares six structured decision-model configurations against multiple baselines to investigate whether providing explicit definitions or decomposing tasks yields practical gains for zero-shot hate speech detection. Experimental results demonstrate that the optimal hosted model approximates the performance of commercial LLMs while reducing inference costs by approximately 97%; however, explicit criteria do not necessarily improve classification accuracy. This work provides empirical evidence supporting low-cost, configurable automated content moderation.