toxicity classification

Designs and trains machine-learning classifiers that detect and label toxic or abusive content (for example offensive, harassing, or hateful material) in user-generated data; this work covers dataset creation and annotation, model selection and training, thresholding and calibration. It also includes evaluating classifier performance, robustness, and fairness across relevant metrics and deployment constraints.

toxicityclassification

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Defining, Understanding, and Detecting Online Toxicity: Challenges and Machine Learning Approaches

Sep 13, 2025
GK
Gautam Kishore Shahi
🏛️ University of Duisburg-Essen | Center for Advanced Internet Studies | University of Bochum

During sensitive periods—such as crises and elections—the detection of online harmful content (e.g., hate speech, offensive language) faces core challenges including conceptual ambiguity and poor generalizability across contexts and languages. Method: This study systematically reviews 140 relevant works to clarify definitional boundaries and data limitations; proposes a novel multilingual, cross-platform toxicity detection paradigm; and introduces a comprehensive benchmark dataset covering 32 languages and high-stakes scenarios—including elections and public health emergencies. Leveraging advanced machine learning and NLP techniques, the study optimizes classification models for enhanced cross-lingual and cross-platform robustness. Contribution/Results: The framework significantly improves accuracy and generalizability in toxic content identification, offering a reusable methodological foundation and empirically grounded guidelines for real-world content moderation practices.

Challenges in automated toxicity detection mechanismsDefining and detecting online toxic contentImproving classification models with cross-platform data

The multi-type, overlapping nature of online hate speech renders conventional binary classification inadequate, motivating the shift toward multi-label classification. Method: We conduct the first systematic review of 46 English-language studies—spanning 28 datasets and 24 models—employing meta-analysis, cross-dataset consistency evaluation, and quantitative assessment of annotation quality (e.g., inter-annotator agreement, IAA). Contribution/Results: We reveal substantial heterogeneity in label taxonomies, dataset sizes, annotation rigor, and evaluation metrics. Key shared challenges include class imbalance, crowdsourcing bias, and sparse minority-label instances. Based on empirical findings, we propose 10 actionable methodological recommendations. We empirically validate the effectiveness of mainstream multi-label architectures—including BERT- and RNN-based models. This work establishes the first academic benchmark and practical guideline for developing robust, comparable, and regulation-compliant multi-label hate speech detection systems.

Analyzing dataset heterogeneity and model evaluation inconsistenciesIdentifying open issues in hate speech classification researchSurveying multi-label hate speech classification models and datasets

ToxiCraft: A Novel Framework for Synthetic Generation of Harmful Information

Sep 23, 2024
ZH
Zheng Hui
🏛️ Microsoft Corporation | Columbia University | Tsinghua University

To address data scarcity and inconsistent labeling criteria for harmful content detection in low-resource settings, this paper proposes ToxiCraft—a framework that generates high-fidelity, diverse toxic texts from minimal seed data. Methodologically, ToxiCraft introduces a novel synthesis paradigm integrating semantic-controllable perturbation with toxicity-aligned distillation, combining prompt-driven generation, adversarial toxicity enhancement, consistency-based filtering, and lightweight discriminator-guided refinement. This design significantly improves model robustness against spurious features and cross-domain generalization. Experiments across multiple benchmarks demonstrate substantial gains in detection accuracy and robustness; generated samples achieve performance on par with human-annotated data, effectively reducing reliance on large-scale manual annotation.

Inconsistent definitions for judging harmful informationLack of data in low-resource harmful content detectionNeed robust models for diverse toxic content classification

Cross-Platform Violence Detection on Social Media: A Dataset and Analysis

May 19, 2025
CC
Celia Chen
🏛️ University of Maryland | Rensselaer Polytechnic Institute

Cross-platform violent content detection is hindered by the scarcity of high-quality, fine-grained annotated datasets—particularly those covering subtypes such as political and sexual violence across multiple platforms. Method: We construct the first large-scale, manually annotated cross-platform violent threat dataset comprising 30,000 instances from Weibo, Twitter, and Reddit, supporting both binary classification and fine-grained multi-subtype recognition. We conduct supervised learning and cross-platform transfer evaluation to assess representational consistency. Contribution/Results: Empirical results demonstrate strong cross-platform consistency in violent content representations: models trained on a single platform achieve high accuracy when tested on others, and performance further improves when training on merged multi-source data. This challenges the “platform-isolated modeling” assumption and validates the semantic transferability of violent content representations. Our dataset and findings provide a critical empirical foundation and methodological validation for robust, generalizable cross-platform content safety governance.

Creating a high-quality dataset for violence classification researchDetecting violent threats across different social media platformsEvaluating cross-platform machine learning models for violence detection

Sensitive Content Classification in Social Media: A Holistic Resource and Evaluation

Nov 29, 2024
DA
Dimosthenis Antypas
🏛️ Cardiff University | University of Mannheim

Existing social media sensitive content detection tools suffer from limited customizability, narrow category coverage—particularly lacking long-tail classes such as drug-related and self-harm content—high privacy risks, and the absence of a unified evaluation benchmark. To address these issues, this work introduces the first high-quality, uniformly annotated dataset covering six sensitive content categories: conflict language, abuse, pornography, drug-related content, self-harm, and spam. We establish standardized protocols for data collection and human annotation. Leveraging this dataset, we supervise fine-tuning of open-source large language models (e.g., LLaMA) and design a comprehensive, multi-dimensional evaluation benchmark. Experimental results demonstrate that our approach consistently outperforms both the LLaMA baseline and the OpenAI API across all six detection tasks, achieving average improvements of 10–15%. Gains are especially pronounced for scarce categories (e.g., drug-related and self-harm content), validating the effectiveness and deployability of open-source LLM fine-tuning for fine-grained sensitive content identification.

Addressing limitations in current moderation tools and datasetsDetecting diverse sensitive content in social media dataImproving accuracy across six key sensitive content categories

Latest Papers

What's happening recently
View more

本文研究了现有仇恨言论检测器对大语言模型生成的仇恨内容的泛化能力,并评估了其随时间和技术进步的稳定性,发现新模型已采取措施防止生成有害内容。

Content ModerationDigital SafetyHate speech

This study addresses the urgent need for effective detection and neutralization of hate speech proliferating on social media. It systematically evaluates the performance of CNN, LSTM, BERT, and their variants in hate speech identification and proposes a novel text transformation method that automatically converts harmful content into semantically preserved neutral expressions. Furthermore, a hybrid model integrating the strengths of multiple architectures is developed, significantly enhancing detection accuracy in specific scenarios. Experimental results demonstrate that BERT-based models achieve superior performance owing to their deep contextual understanding, while the proposed text transformation strategy effectively mitigates the adverse impact of toxic content. The findings validate the feasibility and efficacy of a synergistic framework that jointly performs detection and neutralization.

content moderationhate speechoffensive language

This study addresses the poor performance of existing general-purpose toxicity detection models in real-time gaming chat by identifying a critical gap in high-quality, fine-grained annotated datasets and domain-specific tools through a systematic literature review. To bridge this gap, the authors collaborated with eight League of Legends experts to construct L2DTnH, a fine-grained dataset comprising 1.4k toxic and 13.8k non-toxic messages. Leveraging this dataset, they trained a specialized NLP toxicity detection model and implemented it as a lightweight browser extension that operates locally without reliance on third-party AI services, enabling real-time in-game intervention. Experimental results demonstrate that the proposed model significantly outperforms both general-purpose and state-of-the-art toxicity detectors in gaming contexts and exhibits strong cross-game generalization. The dataset, model, and tool are publicly released.

harassmentNLPonline multiplayer

Adversarially Robust Detection of Harmful Online Content: A Computational Design Science Approach

Dec 19, 2025
YC
Yidong Chai
🏛️ Hefei University of Technology | University of South Florida | University of Georgia | University of Maryland

Automated detection of harmful social media content—such as hate speech, rumors, and extremist text—is vulnerable to adversarial textual perturbations, leading to increased false negatives and poor generalization. To address this, we propose LLM-SGA-ARHOCD: a framework that first leverages large language models to generate and aggregate diverse adversarial samples (LLM-SGA), thereby enhancing attack coverage; it then introduces an Adaptive Robust Hierarchical Online Content Detector (ARHOCD), integrating multi-base model ensembling, Bayesian dynamic weighting, and domain-knowledge-guided collaborative adversarial training. Evaluated on three real-world datasets, our method achieves significant improvements in adversarial robustness (+12.7% on average) and clean-sample accuracy (+3.4% on average), while demonstrating strong cross-attack generalization and high precision. This work establishes a scalable, robust paradigm for secure online content moderation.

Achieve both high generalizability and accuracy in detectionDetect harmful online content robustly against adversarial attacksOvercome limitations of existing adversarial robustness enhancement methods

This work addresses the limitations of existing abuse detection methods, which rely on static models and manual annotations and struggle to handle dynamic, context-sensitive online abusive behaviors. It proposes the first large language model (LLM)-integrated framework spanning the entire lifecycle of abuse detection, systematically encompassing label and feature generation, detection, appeal review, and audit governance. Bridging academic research and industrial practice, the study explores LLMs’ capabilities in contextual reasoning, policy interpretation, and cross-modal understanding. The authors delineate key architectural considerations for each phase, evaluate LLMs’ potential and limitations regarding interpretability, policy alignment, and multimodal fusion, and identify critical challenges—including latency, cost, determinism, adversarial robustness, and fairness—thereby offering a principled direction toward building reliable, accountable, large-scale abuse governance systems.

abuse detectioncontent moderationlarge language models

Hot Scholars

EF

Emilio Ferrara

Professor of Computer Science at the University of Southern California
Human-Centered AISocial ComputingNetwork ScienceAI Safety
AV

Ana Vranić

Social Physics and Complexity (SPAC), LIP, Portugal
statistical physicscomplex systemscomplex networks
AV

Aditya Vashistha

Assistant Professor @ Cornell University
HCIICTDAccessibilityResponsible AI
RD

Ranjie Duan

Alibaba Group
AIAI 安全AI推动共同富裕