Score
Designs and trains machine-learning classifiers that detect and label toxic or abusive content (for example offensive, harassing, or hateful material) in user-generated data; this work covers dataset creation and annotation, model selection and training, thresholding and calibration. It also includes evaluating classifier performance, robustness, and fairness across relevant metrics and deployment constraints.
Online textual abuse—including hate speech and cyberbullying—seriously harms users’ mental health and erodes social trust. While large language models (LLMs) enhance detection capabilities, they may also generate harmful content, exacerbating governance challenges. This study systematically reviews text abuse detection methods in Chinese social media and introduces, for the first time, a “technical–ethical” co-analysis framework. We empirically evaluate leading LLMs across four critical dimensions: detection accuracy, bias, robustness, and risk of generating abusive content. By integrating text classification, psychosocial impact modeling, and adversarial generation analysis, we uncover the dialectical role of LLMs—both mitigating and amplifying online abuse. Our findings provide empirically grounded, actionable insights for safe AI governance, including a phased technical roadmap for responsible deployment and mitigation.
During sensitive periods—such as crises and elections—the detection of online harmful content (e.g., hate speech, offensive language) faces core challenges including conceptual ambiguity and poor generalizability across contexts and languages. Method: This study systematically reviews 140 relevant works to clarify definitional boundaries and data limitations; proposes a novel multilingual, cross-platform toxicity detection paradigm; and introduces a comprehensive benchmark dataset covering 32 languages and high-stakes scenarios—including elections and public health emergencies. Leveraging advanced machine learning and NLP techniques, the study optimizes classification models for enhanced cross-lingual and cross-platform robustness. Contribution/Results: The framework significantly improves accuracy and generalizability in toxic content identification, offering a reusable methodological foundation and empirically grounded guidelines for real-world content moderation practices.
The multi-type, overlapping nature of online hate speech renders conventional binary classification inadequate, motivating the shift toward multi-label classification. Method: We conduct the first systematic review of 46 English-language studies—spanning 28 datasets and 24 models—employing meta-analysis, cross-dataset consistency evaluation, and quantitative assessment of annotation quality (e.g., inter-annotator agreement, IAA). Contribution/Results: We reveal substantial heterogeneity in label taxonomies, dataset sizes, annotation rigor, and evaluation metrics. Key shared challenges include class imbalance, crowdsourcing bias, and sparse minority-label instances. Based on empirical findings, we propose 10 actionable methodological recommendations. We empirically validate the effectiveness of mainstream multi-label architectures—including BERT- and RNN-based models. This work establishes the first academic benchmark and practical guideline for developing robust, comparable, and regulation-compliant multi-label hate speech detection systems.
To address data scarcity and inconsistent labeling criteria for harmful content detection in low-resource settings, this paper proposes ToxiCraft—a framework that generates high-fidelity, diverse toxic texts from minimal seed data. Methodologically, ToxiCraft introduces a novel synthesis paradigm integrating semantic-controllable perturbation with toxicity-aligned distillation, combining prompt-driven generation, adversarial toxicity enhancement, consistency-based filtering, and lightweight discriminator-guided refinement. This design significantly improves model robustness against spurious features and cross-domain generalization. Experiments across multiple benchmarks demonstrate substantial gains in detection accuracy and robustness; generated samples achieve performance on par with human-annotated data, effectively reducing reliance on large-scale manual annotation.
Cross-platform violent content detection is hindered by the scarcity of high-quality, fine-grained annotated datasets—particularly those covering subtypes such as political and sexual violence across multiple platforms. Method: We construct the first large-scale, manually annotated cross-platform violent threat dataset comprising 30,000 instances from Weibo, Twitter, and Reddit, supporting both binary classification and fine-grained multi-subtype recognition. We conduct supervised learning and cross-platform transfer evaluation to assess representational consistency. Contribution/Results: Empirical results demonstrate strong cross-platform consistency in violent content representations: models trained on a single platform achieve high accuracy when tested on others, and performance further improves when training on merged multi-source data. This challenges the “platform-isolated modeling” assumption and validates the semantic transferability of violent content representations. Our dataset and findings provide a critical empirical foundation and methodological validation for robust, generalizable cross-platform content safety governance.
Existing social media sensitive content detection tools suffer from limited customizability, narrow category coverage—particularly lacking long-tail classes such as drug-related and self-harm content—high privacy risks, and the absence of a unified evaluation benchmark. To address these issues, this work introduces the first high-quality, uniformly annotated dataset covering six sensitive content categories: conflict language, abuse, pornography, drug-related content, self-harm, and spam. We establish standardized protocols for data collection and human annotation. Leveraging this dataset, we supervise fine-tuning of open-source large language models (e.g., LLaMA) and design a comprehensive, multi-dimensional evaluation benchmark. Experimental results demonstrate that our approach consistently outperforms both the LLaMA baseline and the OpenAI API across all six detection tasks, achieving average improvements of 10–15%. Gains are especially pronounced for scarce categories (e.g., drug-related and self-harm content), validating the effectiveness and deployability of open-source LLM fine-tuning for fine-grained sensitive content identification.
本文研究了现有仇恨言论检测器对大语言模型生成的仇恨内容的泛化能力,并评估了其随时间和技术进步的稳定性,发现新模型已采取措施防止生成有害内容。
This study addresses the urgent need for effective detection and neutralization of hate speech proliferating on social media. It systematically evaluates the performance of CNN, LSTM, BERT, and their variants in hate speech identification and proposes a novel text transformation method that automatically converts harmful content into semantically preserved neutral expressions. Furthermore, a hybrid model integrating the strengths of multiple architectures is developed, significantly enhancing detection accuracy in specific scenarios. Experimental results demonstrate that BERT-based models achieve superior performance owing to their deep contextual understanding, while the proposed text transformation strategy effectively mitigates the adverse impact of toxic content. The findings validate the feasibility and efficacy of a synergistic framework that jointly performs detection and neutralization.
This study addresses the poor performance of existing general-purpose toxicity detection models in real-time gaming chat by identifying a critical gap in high-quality, fine-grained annotated datasets and domain-specific tools through a systematic literature review. To bridge this gap, the authors collaborated with eight League of Legends experts to construct L2DTnH, a fine-grained dataset comprising 1.4k toxic and 13.8k non-toxic messages. Leveraging this dataset, they trained a specialized NLP toxicity detection model and implemented it as a lightweight browser extension that operates locally without reliance on third-party AI services, enabling real-time in-game intervention. Experimental results demonstrate that the proposed model significantly outperforms both general-purpose and state-of-the-art toxicity detectors in gaming contexts and exhibits strong cross-game generalization. The dataset, model, and tool are publicly released.
Automated detection of harmful social media content—such as hate speech, rumors, and extremist text—is vulnerable to adversarial textual perturbations, leading to increased false negatives and poor generalization. To address this, we propose LLM-SGA-ARHOCD: a framework that first leverages large language models to generate and aggregate diverse adversarial samples (LLM-SGA), thereby enhancing attack coverage; it then introduces an Adaptive Robust Hierarchical Online Content Detector (ARHOCD), integrating multi-base model ensembling, Bayesian dynamic weighting, and domain-knowledge-guided collaborative adversarial training. Evaluated on three real-world datasets, our method achieves significant improvements in adversarial robustness (+12.7% on average) and clean-sample accuracy (+3.4% on average), while demonstrating strong cross-attack generalization and high precision. This work establishes a scalable, robust paradigm for secure online content moderation.
This work addresses the limitations of existing abuse detection methods, which rely on static models and manual annotations and struggle to handle dynamic, context-sensitive online abusive behaviors. It proposes the first large language model (LLM)-integrated framework spanning the entire lifecycle of abuse detection, systematically encompassing label and feature generation, detection, appeal review, and audit governance. Bridging academic research and industrial practice, the study explores LLMs’ capabilities in contextual reasoning, policy interpretation, and cross-modal understanding. The authors delineate key architectural considerations for each phase, evaluate LLMs’ potential and limitations regarding interpretability, policy alignment, and multimodal fusion, and identify critical challenges—including latency, cost, determinism, adversarial robustness, and fairness—thereby offering a principled direction toward building reliable, accountable, large-scale abuse governance systems.