create schema-based annotations

Design and implement schema-based annotation systems for labeling toxicity in text, creating structured label sets, hierarchical categories, and guidelines that capture explicit and implicit toxic behaviors. Produce annotated datasets and mappings to existing taxonomies, plus schema documentation that yields explainable, consistent labels and supports annotation quality control and evaluation of schema-based model predictions.

createschema-basedannotations

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$203K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing toxicity detection methods struggle to effectively identify context-dependent implicit harmful content in multilingual conversations. To address this challenge, this work introduces a multilingual, context-preserving Reddit dataset comprising 125,000 training comments and nearly 3,000 test comments, along with the first systematic framework that jointly models multilingualism, conversational context, and implicit toxicity. The proposed framework employs a hierarchical reasoning architecture integrating context-aware preprocessing, large language model–based automatic annotation, native-speaker human validation, and a hybrid training strategy combining prompt-based learning and fine-tuning to enable fine-grained and interpretable toxicity labeling and evaluation. Experimental results show that while baseline models outperform random guessing, their performance remains substantially suboptimal, underscoring both the difficulty of the task and the necessity of the proposed approach.

contextual understandingimplicit toxicitymultilingual dataset

Rethinking Toxicity Evaluation in Large Language Models: A Multi-Label Perspective

Oct 16, 2025
ZK
Zhiqiang Kou
🏛️ Southeast University | RIKEN Center for Advanced Intelligence Project (AIP) | Qilu University of Technology | The University of Tokyo

Existing LLM toxicity evaluation relies on single-label benchmarks, failing to capture the ambiguity and multidimensional nature of real-world prompts—leading to frequent false negatives and false positives—while fine-grained multi-label human annotation remains prohibitively expensive. Method: We propose a novel multi-label paradigm for toxicity detection, introducing three large-scale, expert-validated multi-label benchmarks—Q-A-MLL, R-A-MLL, and H-X-MLL—covering 15 fine-grained toxicity categories. We theoretically prove the superiority of pseudo-labeling over single-label supervision and leverage public data for cost-effective, high-quality labeling. Contribution/Results: Experiments demonstrate that our approach significantly outperforms strong baselines—including GPT-4o and DeepSeek—in multi-label toxicity classification accuracy and robustness. Our framework establishes a more realistic, scalable, and technically grounded evaluation infrastructure for LLM safety assessment.

Current toxicity detectors fail to capture multi-dimensional toxic contentGathering comprehensive multi-label toxicity annotations is prohibitively expensiveSingle-label benchmarks cause biased evaluations with missed and false detections

Concept-Based Interpretability for Toxicity Detection

Nov 15, 2025
SG
Samarth Garg
🏛️ ABV–IIITM | IIT Jodhpur | IIT Patna

This work addresses the lack of concept-level interpretability and imbalanced concept attribution—leading to misclassifications—in toxic language detection. Methodologically: (1) it treats semantic subtypes (e.g., insult, threat, identity attack) as interpretable concepts, constructs a target lexicon, and proposes a Word–Concept Alignment (WCA) score to quantify each token’s contribution to misclassification via concept gradients (CG); (2) it introduces, for the first time, a delexicalized generative data augmentation strategy to assess model reliance on abstract toxic patterns rather than surface lexical cues. Experiments demonstrate that CG precisely identifies critical toxic tokens and reveal that models over-attribute toxicity to conceptual features even when explicit toxic words are absent—exposing systematic generalization biases toward deep semantic patterns and implicit dependency mechanisms. This establishes a novel paradigm for interpretable toxic language detection grounded in concept-level attribution and causal probing.

Addressing misclassification caused by disproportionate concept attributionAnalyzing model behavior when explicit toxic lexicons are removedDeveloping concept-based interpretability for toxicity detection models

ToxiCraft: A Novel Framework for Synthetic Generation of Harmful Information

Sep 23, 2024
ZH
Zheng Hui
🏛️ Microsoft Corporation | Columbia University | Tsinghua University

To address data scarcity and inconsistent labeling criteria for harmful content detection in low-resource settings, this paper proposes ToxiCraft—a framework that generates high-fidelity, diverse toxic texts from minimal seed data. Methodologically, ToxiCraft introduces a novel synthesis paradigm integrating semantic-controllable perturbation with toxicity-aligned distillation, combining prompt-driven generation, adversarial toxicity enhancement, consistency-based filtering, and lightweight discriminator-guided refinement. This design significantly improves model robustness against spurious features and cross-domain generalization. Experiments across multiple benchmarks demonstrate substantial gains in detection accuracy and robustness; generated samples achieve performance on par with human-annotated data, effectively reducing reliance on large-scale manual annotation.

Inconsistent definitions for judging harmful informationLack of data in low-resource harmful content detectionNeed robust models for diverse toxic content classification

On the Role of Speech Data in Reducing Toxicity Detection Bias

Nov 12, 2024
SJ
Samuel J. Bell
🏛️ Meta | New York University | University College London | McGill University

This work addresses spurious false-positive bias in text toxicity detection—where references to specific demographic groups trigger erroneous toxicity classifications—and presents the first empirical investigation into whether speech modality mitigates this bias. Leveraging the multilingual MuTox dataset, we introduce fine-grained manual demographic annotations and propose a speech-text joint modeling framework. Using fairness metrics—including demographic parity difference (DPD) and equalized odds (EO)—we systematically compare bias in speech versus text classifiers. Results show that speech input reduces group-associated false-positive rates by 23% on average, with especially pronounced gains on ambiguous and contentious samples; moreover, optimizing classifier architecture yields greater bias reduction than improving ASR transcription quality. Our contributions include: (1) empirically demonstrating speech’s bias-mitigation potential; (2) releasing fully annotated data and a multimodal toxicity construction guideline; and (3) establishing a novel paradigm for fair, robust cross-modal content safety detection.

Comparing speech- and text-based toxicity classifiers systematicallyImproving classifiers to reduce group bias in toxicity detectionInvestigating bias reduction in toxicity detection using speech data

Latest Papers

What's happening recently
View more

This study addresses the frequent misclassification of medical terminology and minority-related discourse as harmful content by online moderation systems. Focusing on Bulgarian, the work presents the first toxicity language ontology for the language and introduces a novel fine-grained annotated dataset comprising four categories designed to preserve sensitive yet non-toxic content. By integrating ontology-guided constraints with BERT fine-tuning, the authors train a model on 4,384 manually labeled sentences, achieving a macro-averaged F1 score of 0.89. The proposed approach effectively distinguishes genuinely toxic utterances from critical non-toxic information, offering a deployable solution that significantly enhances both the accuracy and inclusivity of real-world content moderation systems.

BERTBulgarian NLPcontent moderation

This work addresses the challenges posed by the implicit and context-dependent nature of toxic discourse on GitHub, which hinders large-scale annotation and limits the generalizability of existing research. To overcome this, the authors propose a low-cost human-in-the-loop annotation framework that leverages a local small language model to generate toxicity predictions along with interpretable event-category scores. A lightweight random forest verifier then identifies high-risk samples for human review. This approach substantially reduces annotation costs while outperforming confidence-based and multi-model baselines in both efficiency and accuracy. Applying the framework, the study successfully annotated 124,000 GitHub issue and pull request discussions, revising conclusions from prior small-scale studies and systematically uncovering the prevalence, characteristics, and dynamic evolution of toxic behavior in open-source communities.

annotationGitHubhuman-in-the-loop

This work addresses the inconsistent detection and mitigation of toxic content by multilingual large language models across diverse linguistic and cultural contexts. It presents the first systematic synthesis of research on multilingual toxicity handling, proposing a comprehensive framework that encompasses threat modeling, task formulation, detection strategies—such as cross-lingual encoders, translation pipelines, and representation probing—and mitigation approaches, including data filtering, alignment tuning, decoding controls, and multilingual safeguards. The study identifies core challenges such as uneven language coverage and culturally contingent definitions of harm, while highlighting critical issues like fragmented evaluation protocols and the unintended suppression of legitimate expression. By elucidating these dimensions, the paper establishes a theoretical foundation and practical roadmap for achieving cross-lingual safety alignment in multilingual language models.

cultural contextharm mitigationmultilingual language models

This work addresses the challenge of aligning language models with users’ subjective sensitivities to harmful content without relying on global alignment standards. It systematically evaluates training-free, inference-stage interventions across three phases—pre-decoding, during decoding, and post-decoding—to enable personalized control over toxicity sensitivity. The study presents the first comprehensive comparison of diverse training-free techniques, including prompt modulation, token/logit/representation manipulation, and candidate re-ranking, in the context of personalized toxicity alignment. Through this analysis, it reveals inherent trade-offs among alignment accuracy, degree of personalization, and linguistic quality. Experimental results on the PRISM dataset demonstrate that the proposed approaches reduce alignment error by 28%–47%, establishing the feasibility and effectiveness of achieving personalized alignment without model retraining.

alignmentlanguage modelspersonalization

Scientific process descriptions are often embedded in unstructured text, hindering reproducibility, comparison, and automation. To address this challenge, this work presents the first cross-disciplinary, expert-driven repository of structured scientific process schemas, encompassing 16 expert-annotated patterns across five domains. Through a human-in-the-loop workflow, candidate schemas generated by large language models were iteratively refined via domain expert feedback, yielding reusable fields such as inputs, outputs, steps, and parameters. The resulting schemas are formalized in both JSON Schema and SHACL formats and accompanied by an integrated toolchain. The project also releases a comprehensive dataset—including schemas, intermediate artifacts, review records, and analysis scripts—to support knowledge graph construction, semantic publishing, and cross-study comparison.

heterogeneous datareproducibilityschema

Hot Scholars

GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
AJ

Adam Jatowt

Professor at Univ. of Innsbruck (previously Kyoto Univ.)
question answeringlarge language modelsinformation retrievalRAG
JS

Jinwook Seo

Department of Computer Science and Engineering, Seoul National University
Human-Computer InteractionInformation VisualizationVisual AnalyticsExplainable AI
MK

Mohsinul Kabir

PhD Candidate at the University of Manchester
NLPHCIAI
CH

Conghui He

Shanghai AI Laboratory
Data-centric AILLMDocument Intelligence