Score
Design and implement schema-based annotation systems for labeling toxicity in text, creating structured label sets, hierarchical categories, and guidelines that capture explicit and implicit toxic behaviors. Produce annotated datasets and mappings to existing taxonomies, plus schema documentation that yields explainable, consistent labels and supports annotation quality control and evaluation of schema-based model predictions.
Existing toxicity detection methods struggle to effectively identify context-dependent implicit harmful content in multilingual conversations. To address this challenge, this work introduces a multilingual, context-preserving Reddit dataset comprising 125,000 training comments and nearly 3,000 test comments, along with the first systematic framework that jointly models multilingualism, conversational context, and implicit toxicity. The proposed framework employs a hierarchical reasoning architecture integrating context-aware preprocessing, large language model–based automatic annotation, native-speaker human validation, and a hybrid training strategy combining prompt-based learning and fine-tuning to enable fine-grained and interpretable toxicity labeling and evaluation. Experimental results show that while baseline models outperform random guessing, their performance remains substantially suboptimal, underscoring both the difficulty of the task and the necessity of the proposed approach.
Existing LLM toxicity evaluation relies on single-label benchmarks, failing to capture the ambiguity and multidimensional nature of real-world prompts—leading to frequent false negatives and false positives—while fine-grained multi-label human annotation remains prohibitively expensive. Method: We propose a novel multi-label paradigm for toxicity detection, introducing three large-scale, expert-validated multi-label benchmarks—Q-A-MLL, R-A-MLL, and H-X-MLL—covering 15 fine-grained toxicity categories. We theoretically prove the superiority of pseudo-labeling over single-label supervision and leverage public data for cost-effective, high-quality labeling. Contribution/Results: Experiments demonstrate that our approach significantly outperforms strong baselines—including GPT-4o and DeepSeek—in multi-label toxicity classification accuracy and robustness. Our framework establishes a more realistic, scalable, and technically grounded evaluation infrastructure for LLM safety assessment.
This work addresses the lack of concept-level interpretability and imbalanced concept attribution—leading to misclassifications—in toxic language detection. Methodologically: (1) it treats semantic subtypes (e.g., insult, threat, identity attack) as interpretable concepts, constructs a target lexicon, and proposes a Word–Concept Alignment (WCA) score to quantify each token’s contribution to misclassification via concept gradients (CG); (2) it introduces, for the first time, a delexicalized generative data augmentation strategy to assess model reliance on abstract toxic patterns rather than surface lexical cues. Experiments demonstrate that CG precisely identifies critical toxic tokens and reveal that models over-attribute toxicity to conceptual features even when explicit toxic words are absent—exposing systematic generalization biases toward deep semantic patterns and implicit dependency mechanisms. This establishes a novel paradigm for interpretable toxic language detection grounded in concept-level attribution and causal probing.
To address data scarcity and inconsistent labeling criteria for harmful content detection in low-resource settings, this paper proposes ToxiCraft—a framework that generates high-fidelity, diverse toxic texts from minimal seed data. Methodologically, ToxiCraft introduces a novel synthesis paradigm integrating semantic-controllable perturbation with toxicity-aligned distillation, combining prompt-driven generation, adversarial toxicity enhancement, consistency-based filtering, and lightweight discriminator-guided refinement. This design significantly improves model robustness against spurious features and cross-domain generalization. Experiments across multiple benchmarks demonstrate substantial gains in detection accuracy and robustness; generated samples achieve performance on par with human-annotated data, effectively reducing reliance on large-scale manual annotation.
This work addresses spurious false-positive bias in text toxicity detection—where references to specific demographic groups trigger erroneous toxicity classifications—and presents the first empirical investigation into whether speech modality mitigates this bias. Leveraging the multilingual MuTox dataset, we introduce fine-grained manual demographic annotations and propose a speech-text joint modeling framework. Using fairness metrics—including demographic parity difference (DPD) and equalized odds (EO)—we systematically compare bias in speech versus text classifiers. Results show that speech input reduces group-associated false-positive rates by 23% on average, with especially pronounced gains on ambiguous and contentious samples; moreover, optimizing classifier architecture yields greater bias reduction than improving ASR transcription quality. Our contributions include: (1) empirically demonstrating speech’s bias-mitigation potential; (2) releasing fully annotated data and a multimodal toxicity construction guideline; and (3) establishing a novel paradigm for fair, robust cross-modal content safety detection.
This study addresses the frequent misclassification of medical terminology and minority-related discourse as harmful content by online moderation systems. Focusing on Bulgarian, the work presents the first toxicity language ontology for the language and introduces a novel fine-grained annotated dataset comprising four categories designed to preserve sensitive yet non-toxic content. By integrating ontology-guided constraints with BERT fine-tuning, the authors train a model on 4,384 manually labeled sentences, achieving a macro-averaged F1 score of 0.89. The proposed approach effectively distinguishes genuinely toxic utterances from critical non-toxic information, offering a deployable solution that significantly enhances both the accuracy and inclusivity of real-world content moderation systems.
This work addresses the challenges posed by the implicit and context-dependent nature of toxic discourse on GitHub, which hinders large-scale annotation and limits the generalizability of existing research. To overcome this, the authors propose a low-cost human-in-the-loop annotation framework that leverages a local small language model to generate toxicity predictions along with interpretable event-category scores. A lightweight random forest verifier then identifies high-risk samples for human review. This approach substantially reduces annotation costs while outperforming confidence-based and multi-model baselines in both efficiency and accuracy. Applying the framework, the study successfully annotated 124,000 GitHub issue and pull request discussions, revising conclusions from prior small-scale studies and systematically uncovering the prevalence, characteristics, and dynamic evolution of toxic behavior in open-source communities.
This work addresses the inconsistent detection and mitigation of toxic content by multilingual large language models across diverse linguistic and cultural contexts. It presents the first systematic synthesis of research on multilingual toxicity handling, proposing a comprehensive framework that encompasses threat modeling, task formulation, detection strategies—such as cross-lingual encoders, translation pipelines, and representation probing—and mitigation approaches, including data filtering, alignment tuning, decoding controls, and multilingual safeguards. The study identifies core challenges such as uneven language coverage and culturally contingent definitions of harm, while highlighting critical issues like fragmented evaluation protocols and the unintended suppression of legitimate expression. By elucidating these dimensions, the paper establishes a theoretical foundation and practical roadmap for achieving cross-lingual safety alignment in multilingual language models.
This work addresses the challenge of aligning language models with users’ subjective sensitivities to harmful content without relying on global alignment standards. It systematically evaluates training-free, inference-stage interventions across three phases—pre-decoding, during decoding, and post-decoding—to enable personalized control over toxicity sensitivity. The study presents the first comprehensive comparison of diverse training-free techniques, including prompt modulation, token/logit/representation manipulation, and candidate re-ranking, in the context of personalized toxicity alignment. Through this analysis, it reveals inherent trade-offs among alignment accuracy, degree of personalization, and linguistic quality. Experimental results on the PRISM dataset demonstrate that the proposed approaches reduce alignment error by 28%–47%, establishing the feasibility and effectiveness of achieving personalized alignment without model retraining.
Scientific process descriptions are often embedded in unstructured text, hindering reproducibility, comparison, and automation. To address this challenge, this work presents the first cross-disciplinary, expert-driven repository of structured scientific process schemas, encompassing 16 expert-annotated patterns across five domains. Through a human-in-the-loop workflow, candidate schemas generated by large language models were iteratively refined via domain expert feedback, yielding reusable fields such as inputs, outputs, steps, and parameters. The resulting schemas are formalized in both JSON Schema and SHACL formats and accompanied by an integrated toolchain. The project also releases a comprehensive dataset—including schemas, intermediate artifacts, review records, and analysis scripts—to support knowledge graph construction, semantic publishing, and cross-study comparison.