multilingual data collection

Designs and builds contextual multilingual datasets and corpora for safety, toxicity, and moderation tasks, including bilingual and translated collections, annotated harm/toxicity datasets, and curated evaluation benchmarks that preserve conversational thread context, entity spans, and relation labels. Implements multilingual data sourcing, preprocessing, curation, annotation, and generation—controlling code‑mixing and linguistic diversity—to produce representative training and evaluation data for both high‑ and low‑resource languages.

multilingualdatacollection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.29
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$201K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the inconsistent detection and mitigation of toxic content by multilingual large language models across diverse linguistic and cultural contexts. It presents the first systematic synthesis of research on multilingual toxicity handling, proposing a comprehensive framework that encompasses threat modeling, task formulation, detection strategies—such as cross-lingual encoders, translation pipelines, and representation probing—and mitigation approaches, including data filtering, alignment tuning, decoding controls, and multilingual safeguards. The study identifies core challenges such as uneven language coverage and culturally contingent definitions of harm, while highlighting critical issues like fragmented evaluation protocols and the unintended suppression of legitimate expression. By elucidating these dimensions, the paper establishes a theoretical foundation and practical roadmap for achieving cross-lingual safety alignment in multilingual language models.

cultural contextharm mitigationmultilingual language models

Existing toxicity detection methods struggle to effectively identify context-dependent implicit harmful content in multilingual conversations. To address this challenge, this work introduces a multilingual, context-preserving Reddit dataset comprising 125,000 training comments and nearly 3,000 test comments, along with the first systematic framework that jointly models multilingualism, conversational context, and implicit toxicity. The proposed framework employs a hierarchical reasoning architecture integrating context-aware preprocessing, large language model–based automatic annotation, native-speaker human validation, and a hybrid training strategy combining prompt-based learning and fine-tuning to enable fine-grained and interpretable toxicity labeling and evaluation. Experimental results show that while baseline models outperform random guessing, their performance remains substantially suboptimal, underscoring both the difficulty of the task and the necessity of the proposed approach.

contextual understandingimplicit toxicitymultilingual dataset

Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation

Jul 16, 2025
ZG
Ziyu Ge
🏛️ Singapore University of Technology and Design | GovTech

Translating toxic content embedded in low-resource Singlish—characterized by slang, code-switching, and culturally grounded harmful expressions—poses significant challenges: existing systems fail to preserve sociolinguistic nuance and toxicity, while suffering from scarce parallel data and insufficient toxicity awareness. Method: We propose a two-stage toxicity-aware translation framework. It integrates human-verified few-shot prompting with model-prompt co-optimization to enable controlled retention of toxic expressions under extreme scarcity of safety-annotated data. Toxicity preservation is rigorously validated via forward/back-translation semantic similarity scoring, expert curation, and cross-LLM comparative verification. Results: Experiments demonstrate substantial improvements in preserving sociolinguistic detail and toxic register, with quantitative metrics and human evaluation confirming effectiveness, robustness, and reproducibility. This work establishes a novel paradigm for socioculturally informed safety governance of multilingual large language models.

Address scarce data and safety filters in translationPreserve sociolinguistic nuance in multicultural content moderationTranslate low-resource Singlish with slang and toxicity

Translate, then Detect: Leveraging Machine Translation for Cross-Lingual Toxicity Classification

Sep 17, 2025
SJ
Samuel J. Bell
🏛️ FAIR at Meta | University College London | University of the Basque Country (UPV/EHU)

Multilingual toxicity detection faces significant challenges due to the scarcity of annotated data for low-resource languages. This paper systematically investigates machine translation (MT)-enabled cross-lingual detection, comparing the “translate-then-classify” paradigm using monolingual classifiers against direct zero-shot inference with multilingual large language models (MLLMs). We further propose an MT-specific fine-tuning strategy tailored to toxicity detection to reduce rejection rates. Experiments across 16 languages show that translate-then-classify outperforms out-of-distribution MLLMs on 81.3% of languages—and achieves statistically significant gains on 6 out of 7 low-resource languages. Our analysis identifies MT quality and underlying language resource availability as key performance determinants. Results validate that lightweight MT coupled with conventional classifiers offers both effectiveness and practicality in low-resource settings, establishing a scalable new paradigm for multilingual content safety under resource constraints.

Assessing performance across resource levels and machine translation qualityComparing translation-based versus language-specific classification pipelinesEvaluating machine translation for cross-lingual toxicity classification

SynthDetoxM: Modern LLMs are Few-Shot Parallel Detoxification Data Annotators

Feb 10, 2025
DM
Daniil Moskovskiy
🏛️ AIRI | Skoltech | Sber AI | ISP RAS | Research Center for Trusted AI

Multilingual text detoxification is hindered by the scarcity of high-quality parallel annotated data, especially for low-resource languages. To address this, we propose the first synthetic data generation framework for multilingual detoxification: leveraging few-shot prompting, we orchestrate nine open-source large language models to generate detoxified parallel sentence pairs in German, French, Spanish, and Russian; these are rigorously filtered via multi-source toxicity scoring and human verification, yielding SynthDetoxM—a 16k-instance benchmark of high-fidelity parallel sentences, the first systematically validated multilingual synthetic detoxification dataset. Models trained on SynthDetoxM significantly outperform those trained on the real-world multilingual dataset MultiParaDetox under data-constrained settings and surpass all baseline LMs in few-shot evaluation across languages. This demonstrates both the efficacy and scalability of synthetic-data-driven detoxification modeling.

Address multilingual text detoxification data scarcityEnhance model performance with few-shot LLMsIntroduce synthetic multilingual detoxification dataset

Latest Papers

What's happening recently
View more

This study addresses coverage gaps and quality deficiencies in monolingual dimensions of multilingual AI safety benchmarks by proposing a reusable slice-level auditing methodology. Through cross-tier empirical comparisons and controlled experiments across 21 resources, we systematically reveal structural deficits in non-English data regarding provenance, annotation, and harm taxonomy. The research demonstrates that these disparities are quantifiably remediable and establishes a causal link between data-sparse regions and asymmetric model robustness against jailbreaking. By filling critical coverage voids in categories such as self-harm, this work renders multilingual safety claims empirically verifiable and provides methodological foundations for constructing high-quality multilingual safety datasets.

AI SafetyDataset GapsJailbreak Robustness

This study addresses the pervasive lack of cultural competence in contemporary multilingual NLP models, which often fail to accurately interpret expressions deeply rooted in specific cultural contexts despite broad language coverage. Synthesizing insights from over 50 studies published between 2020 and 2026, the work advocates a paradigm shift from isolated language processing toward modeling the “communicative ecology,” integrating institutional norms, cultural scripts, and community practices as essential contextual dimensions. Through culturally aware evaluation benchmarks (e.g., Global-MMLU, CulturalBench), multimodal grounding of local knowledge, community-coconstructed datasets, and cultural alignment techniques, the research demonstrates that insufficient training data coverage is not the sole bottleneck—language choice, tokenization strategies, and translation benchmark design are equally critical. The paper calls for layered cultural evaluation frameworks and participatory alignment approaches to advance fair, inclusive, and culturally grounded NLP systems.

community-grounded evaluationcross-lingual transfercultural competence

This study addresses the frequent misclassification of medical terminology and minority-related discourse as harmful content by online moderation systems. Focusing on Bulgarian, the work presents the first toxicity language ontology for the language and introduces a novel fine-grained annotated dataset comprising four categories designed to preserve sensitive yet non-toxic content. By integrating ontology-guided constraints with BERT fine-tuning, the authors train a model on 4,384 manually labeled sentences, achieving a macro-averaged F1 score of 0.89. The proposed approach effectively distinguishes genuinely toxic utterances from critical non-toxic information, offering a deployable solution that significantly enhances both the accuracy and inclusivity of real-world content moderation systems.

BERTBulgarian NLPcontent moderation

This study addresses the performance degradation commonly observed in multilingual large language models, which stems from imbalanced data distributions and the so-called “curse of multilinguality.” The authors identify the root cause as remediable corpus quality issues and propose a language-specific data curation and balancing strategy. By integrating multilingual quality evaluation with an efficient training mixture methodology, they optimize the composition of a 20-trillion-token corpus. Models trained on this refined dataset—specifically 3B and 8B parameter variants—achieve state-of-the-art multilingual performance while using 4–10 times fewer FLOPs than competing approaches. Furthermore, the curated corpus significantly enhances the multilingual scaling efficiency of Trinity Large (400B), demonstrating its effectiveness in improving both model performance and training efficiency across diverse languages.

curse of multilingualitydata curationmultilingual interference

This study addresses the critical gap in safety evaluation resources for large language models (LLMs) in non-English languages, particularly German—a high-resource language—and Bulgarian—a low-resource language—where existing benchmarks inadequately capture risks related to harmful content generation within specific sociocultural and legal contexts. To bridge this gap, the authors introduce the first regionally grounded bilingual safety evaluation dataset covering both languages, constructed through carefully curated and human-crafted adversarial prompts spanning multiple culturally sensitive topics. Systematic evaluations using both multilingual and monolingual LLMs reveal significant cross-lingual disparities in safety behaviors. This work underscores the necessity of localized, region-specific benchmarks for the responsible deployment of LLMs and fills a crucial void in safety assessment for major non-English languages.

BulgarianGermanLLM safety

Hot Scholars

SH

Shamsuddeen Hassan Muhammad

Bayero University, Kano, & Google DeepMind Academic Fellow at Imperial College London
Natural Language ProcessingSentiment AnalysisAfricaNLPLow-resource NLP
AF

Alham Fikri Aji

MBZUAI, Monash Indonesia
MultilingualityLow-resource NLPLanguage ModelingMachine Translation
PN

Preslav Nakov

Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
Computational LinguisticsLarge Language ModelsFact-checkingFake News
DI

David Ifeoluwa Adelani

McGill University and Mila - Quebec AI Institute and Canada CIFAR AI Chair
Natural language processingMultilingualityMultilingual NLPAfricaNLP
IA

Idris Abdulmumin

Postdoctoral Fellow, DSFSI, University of Pretoria
Machine TranslationNeural Machine TranslationNatural Language ProcessingInternet Technology