Score
Designs and builds contextual multilingual datasets and corpora for safety, toxicity, and moderation tasks, including bilingual and translated collections, annotated harm/toxicity datasets, and curated evaluation benchmarks that preserve conversational thread context, entity spans, and relation labels. Implements multilingual data sourcing, preprocessing, curation, annotation, and generation—controlling code‑mixing and linguistic diversity—to produce representative training and evaluation data for both high‑ and low‑resource languages.
This work addresses the inconsistent detection and mitigation of toxic content by multilingual large language models across diverse linguistic and cultural contexts. It presents the first systematic synthesis of research on multilingual toxicity handling, proposing a comprehensive framework that encompasses threat modeling, task formulation, detection strategies—such as cross-lingual encoders, translation pipelines, and representation probing—and mitigation approaches, including data filtering, alignment tuning, decoding controls, and multilingual safeguards. The study identifies core challenges such as uneven language coverage and culturally contingent definitions of harm, while highlighting critical issues like fragmented evaluation protocols and the unintended suppression of legitimate expression. By elucidating these dimensions, the paper establishes a theoretical foundation and practical roadmap for achieving cross-lingual safety alignment in multilingual language models.
Existing toxicity detection methods struggle to effectively identify context-dependent implicit harmful content in multilingual conversations. To address this challenge, this work introduces a multilingual, context-preserving Reddit dataset comprising 125,000 training comments and nearly 3,000 test comments, along with the first systematic framework that jointly models multilingualism, conversational context, and implicit toxicity. The proposed framework employs a hierarchical reasoning architecture integrating context-aware preprocessing, large language model–based automatic annotation, native-speaker human validation, and a hybrid training strategy combining prompt-based learning and fine-tuning to enable fine-grained and interpretable toxicity labeling and evaluation. Experimental results show that while baseline models outperform random guessing, their performance remains substantially suboptimal, underscoring both the difficulty of the task and the necessity of the proposed approach.
Translating toxic content embedded in low-resource Singlish—characterized by slang, code-switching, and culturally grounded harmful expressions—poses significant challenges: existing systems fail to preserve sociolinguistic nuance and toxicity, while suffering from scarce parallel data and insufficient toxicity awareness. Method: We propose a two-stage toxicity-aware translation framework. It integrates human-verified few-shot prompting with model-prompt co-optimization to enable controlled retention of toxic expressions under extreme scarcity of safety-annotated data. Toxicity preservation is rigorously validated via forward/back-translation semantic similarity scoring, expert curation, and cross-LLM comparative verification. Results: Experiments demonstrate substantial improvements in preserving sociolinguistic detail and toxic register, with quantitative metrics and human evaluation confirming effectiveness, robustness, and reproducibility. This work establishes a novel paradigm for socioculturally informed safety governance of multilingual large language models.
Multilingual toxicity detection faces significant challenges due to the scarcity of annotated data for low-resource languages. This paper systematically investigates machine translation (MT)-enabled cross-lingual detection, comparing the “translate-then-classify” paradigm using monolingual classifiers against direct zero-shot inference with multilingual large language models (MLLMs). We further propose an MT-specific fine-tuning strategy tailored to toxicity detection to reduce rejection rates. Experiments across 16 languages show that translate-then-classify outperforms out-of-distribution MLLMs on 81.3% of languages—and achieves statistically significant gains on 6 out of 7 low-resource languages. Our analysis identifies MT quality and underlying language resource availability as key performance determinants. Results validate that lightweight MT coupled with conventional classifiers offers both effectiveness and practicality in low-resource settings, establishing a scalable new paradigm for multilingual content safety under resource constraints.
Multilingual text detoxification is hindered by the scarcity of high-quality parallel annotated data, especially for low-resource languages. To address this, we propose the first synthetic data generation framework for multilingual detoxification: leveraging few-shot prompting, we orchestrate nine open-source large language models to generate detoxified parallel sentence pairs in German, French, Spanish, and Russian; these are rigorously filtered via multi-source toxicity scoring and human verification, yielding SynthDetoxM—a 16k-instance benchmark of high-fidelity parallel sentences, the first systematically validated multilingual synthetic detoxification dataset. Models trained on SynthDetoxM significantly outperform those trained on the real-world multilingual dataset MultiParaDetox under data-constrained settings and surpass all baseline LMs in few-shot evaluation across languages. This demonstrates both the efficacy and scalability of synthetic-data-driven detoxification modeling.
This study addresses coverage gaps and quality deficiencies in monolingual dimensions of multilingual AI safety benchmarks by proposing a reusable slice-level auditing methodology. Through cross-tier empirical comparisons and controlled experiments across 21 resources, we systematically reveal structural deficits in non-English data regarding provenance, annotation, and harm taxonomy. The research demonstrates that these disparities are quantifiably remediable and establishes a causal link between data-sparse regions and asymmetric model robustness against jailbreaking. By filling critical coverage voids in categories such as self-harm, this work renders multilingual safety claims empirically verifiable and provides methodological foundations for constructing high-quality multilingual safety datasets.
This study addresses the pervasive lack of cultural competence in contemporary multilingual NLP models, which often fail to accurately interpret expressions deeply rooted in specific cultural contexts despite broad language coverage. Synthesizing insights from over 50 studies published between 2020 and 2026, the work advocates a paradigm shift from isolated language processing toward modeling the “communicative ecology,” integrating institutional norms, cultural scripts, and community practices as essential contextual dimensions. Through culturally aware evaluation benchmarks (e.g., Global-MMLU, CulturalBench), multimodal grounding of local knowledge, community-coconstructed datasets, and cultural alignment techniques, the research demonstrates that insufficient training data coverage is not the sole bottleneck—language choice, tokenization strategies, and translation benchmark design are equally critical. The paper calls for layered cultural evaluation frameworks and participatory alignment approaches to advance fair, inclusive, and culturally grounded NLP systems.
This study addresses the frequent misclassification of medical terminology and minority-related discourse as harmful content by online moderation systems. Focusing on Bulgarian, the work presents the first toxicity language ontology for the language and introduces a novel fine-grained annotated dataset comprising four categories designed to preserve sensitive yet non-toxic content. By integrating ontology-guided constraints with BERT fine-tuning, the authors train a model on 4,384 manually labeled sentences, achieving a macro-averaged F1 score of 0.89. The proposed approach effectively distinguishes genuinely toxic utterances from critical non-toxic information, offering a deployable solution that significantly enhances both the accuracy and inclusivity of real-world content moderation systems.
This study addresses the performance degradation commonly observed in multilingual large language models, which stems from imbalanced data distributions and the so-called “curse of multilinguality.” The authors identify the root cause as remediable corpus quality issues and propose a language-specific data curation and balancing strategy. By integrating multilingual quality evaluation with an efficient training mixture methodology, they optimize the composition of a 20-trillion-token corpus. Models trained on this refined dataset—specifically 3B and 8B parameter variants—achieve state-of-the-art multilingual performance while using 4–10 times fewer FLOPs than competing approaches. Furthermore, the curated corpus significantly enhances the multilingual scaling efficiency of Trinity Large (400B), demonstrating its effectiveness in improving both model performance and training efficiency across diverse languages.
This study addresses the critical gap in safety evaluation resources for large language models (LLMs) in non-English languages, particularly German—a high-resource language—and Bulgarian—a low-resource language—where existing benchmarks inadequately capture risks related to harmful content generation within specific sociocultural and legal contexts. To bridge this gap, the authors introduce the first regionally grounded bilingual safety evaluation dataset covering both languages, constructed through carefully curated and human-crafted adversarial prompts spanning multiple culturally sensitive topics. Systematic evaluations using both multilingual and monolingual LLMs reveal significant cross-lingual disparities in safety behaviors. This work underscores the necessity of localized, region-specific benchmarks for the responsible deployment of LLMs and fills a crucial void in safety assessment for major non-English languages.