language identification

Automatically detecting the language of text data and curating multilingual corpora, including preprocessing and deduplication steps to assemble high-quality datasets for training.

languageidentification

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

HPLT~3.0: Very Large-Scale Multilingual Resources for LLM and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models

Nov 02, 2025
SO
Stephan Oepen
🏛️ University of Oslo | University of Helsinki | Prompsit Language Engineering | The Common Crawl Foundation | Edinburgh University | Charles University | University of Turku

To address the scarcity of high-quality, large-scale, fine-grained annotated data for multilingual large language models and machine translation research, this work introduces the first open-source, large-scale multilingual text data construction framework. It supports nearly 200 languages and scales to 30 trillion tokens, integrating end-to-end techniques including web page cleaning, noise-robust language identification, exact and fuzzy deduplication, PII detection, registry label annotation, text quality scoring, and synthetic parallel corpus generation. We propose the first native-task-based multilingual evaluation framework, releasing standardized benchmarks across nine languages and an automated assessment pipeline. The resulting dataset constitutes the largest publicly available multilingual pretraining corpus to date. Using it, we train 57 monolingual encoder-decoder models and multiple GPT-style monolingual models. All tooling—including data processing pipelines, evaluation benchmarks, and pretrained model families—is fully open-sourced.

Creating massive open multilingual datasets for LLM training and machine translationDeveloping comprehensive evaluation benchmarks for multilingual language model assessmentProviding high-quality monolingual and bilingual data with rich annotations and filtering

This study addresses the critical challenges in Lombard, a low-resource language, where existing NLP corpora suffer from widespread mislabeling, templated content, noise, and severe dialectal imbalance. For the first time, the authors conduct a systematic manual audit, integrating language identification validation, orthographic analysis, and dialect classification to assess linguistic authenticity and regional representativeness. Their findings reveal that mainstream datasets contain an extremely low proportion of genuine Lombard texts, exhibit inconsistent orthography, and display a pronounced representational bias favoring Western dialects while marginalizing Eastern varieties. The work underscores the necessity of moving beyond quantity-driven data collection toward community-informed, dialect-sensitive strategies for building high-quality linguistic resources, offering crucial methodological guidance for corpus development in other low-resource languages.

corpus auditinglanguage identificationLombard

This study addresses the performance degradation commonly observed in multilingual large language models, which stems from imbalanced data distributions and the so-called “curse of multilinguality.” The authors identify the root cause as remediable corpus quality issues and propose a language-specific data curation and balancing strategy. By integrating multilingual quality evaluation with an efficient training mixture methodology, they optimize the composition of a 20-trillion-token corpus. Models trained on this refined dataset—specifically 3B and 8B parameter variants—achieve state-of-the-art multilingual performance while using 4–10 times fewer FLOPs than competing approaches. Furthermore, the curated corpus significantly enhances the multilingual scaling efficiency of Trinity Large (400B), demonstrating its effectiveness in improving both model performance and training efficiency across diverse languages.

curse of multilingualitydata curationmultilingual interference

To address the challenges of poor generalizability in preprocessing pipelines and uneven data quality in multilingual large language model (LLM) training, this work introduces the first scalable, automated multilingual pretraining data processing framework—supporting up to one thousand languages. Built upon Common Crawl, the framework integrates language-aware efficient filtering, cross-lingual deduplication, and joint optimization of duplication rate and quality for data rebalancing. It employs end-to-end evaluation via ablation studies guided by multilingual downstream tasks. We release FineWeb2, a 5-billion-document, 20-TB multilingual dataset covering nine languages. Experiments demonstrate substantial improvements in non-English LLM performance across multiple benchmarks. This work establishes a systematic, reproducible infrastructure for high-quality multilingual foundation model training.

Adapting pre-training data processing for multilingual LLMsCreating clean diverse datasets for non-English languagesRebalancing datasets considering duplication and quality

WanJuanSiLu: A High-Quality Open-Source Webtext Dataset for Low-Resource Languages

Jan 24, 2025
JY
Jia Yu
🏛️ Shanghai Artificial Intelligence Laboratory

Low-resource languages suffer from a critical scarcity of high-quality multilingual textual data, severely constraining the development of large language models. To address this, we propose the first systematic, open-source web corpus construction framework specifically designed for low-resource languages. Our framework integrates multi-stage collaborative processing: adaptive web page extraction, language-aware cleaning, cross-document semantic deduplication, fine-grained safety filtering, multidimensional quality assessment, and topic-consistency classification—ensuring both linguistic diversity and enhanced data security and reliability. We publicly release high-quality corpora covering five low-resource languages. Empirical evaluations demonstrate superior data quality, safety, and usability compared to existing benchmarks. The datasets and code are fully open-sourced on OpenDataLab and GitHub, providing a robust, scalable foundation for multilingual large model training and research.

Multi-language DatasetsResource-poor LanguagesText Data Quality

Latest Papers

What's happening recently
View more

AraMix: Recycling, Refiltering, and Deduplicating to Deliver the Largest Arabic Pretraining Corpus

Dec 21, 2025
SA
Sultan Alrashed
🏛️ King Abdullah University of Science and Technology (KAUST)

Arabic pretraining corpora suffer from severe redundancy (nearly 60% token-level duplication) and heterogeneous quality. To address this, we propose a “data reuse over new crawling” paradigm, systematically integrating seven existing public Arabic web datasets. Our pipeline applies Arabic-specific quality filtering, MinHash-based deduplication at both document and sentence levels, multi-source fusion, and metadata alignment. The resulting corpus—currently the largest publicly available, deeply deduplicated Arabic dataset—comprises 178 billion tokens across 179 million documents. Empirical evaluation demonstrates substantial improvements in downstream model training efficiency and generalization performance. This corpus has become the de facto standard training data for multiple open-source Arabic large language models, establishing a new principle in Arabic NLP: rigorous, quality-driven data curation takes precedence over mere scale expansion.

Constructing a large, high-quality Arabic pretraining corpusOptimizing data curation over new web scraping for low-resource languagesReducing redundancy and duplicates in existing Arabic datasets

Improving Romanian LLM Pretraining Data using Diversity and Quality Filtering

Nov 02, 2025
VN
Vlad Negoita
🏛️ National University of Science and Technology POLITEHNICA Bucharest

Pretraining data for Romanian is scarce, heterogeneous in quality, and insufficiently diverse in topical coverage. Method: This paper proposes a lightweight multi-task filtering framework that jointly evaluates educational value, performs topic modeling, and analyzes syntactic and formatting diversity to conduct hierarchical quality screening of LLM-generated annotated texts. Unlike conventional unidimensional data cleaning, the framework systematically characterizes cross-lingual disparities between Romanian and English pretraining corpora across topic distribution, pedagogical relevance, and structural diversity. Contribution/Results: Empirical evaluation demonstrates that models trained on the filtered Romanian corpus achieve substantial performance gains on downstream tasks—including ROBUST and RONEC—validating the efficacy of structured, domain-aware data curation for low-resource language modeling. The work establishes a principled paradigm for small-language pretraining data construction, highlighting its critical role in enhancing model capabilities.

Analyzing Romanian pretraining corpora characteristics and English comparisonsDeveloping multi-level filtering methods for high-quality Romanian datasetsImproving Romanian LLM performance through diversity and quality filtering

Multilingual corpora for the study of new concepts in the social sciences and humanities:

Dec 08, 2025
RK
Revekka Kyriakoglou
🏛️ LIASD | Université Paris 8

This study addresses the scarcity of multilingual corpora supporting emerging concepts—such as “non-technological innovation”—in the humanities and social sciences (HSS). To tackle this, we propose a hybrid multilingual corpus construction methodology that integrates corporate websites and annual reports, combining automatic language identification, domain-adapted content filtering, relevant paragraph extraction, expert-lexicon-driven contextual block identification, thematic annotation, and enriched structured metadata. Our key contribution is the first systematic construction of a high-quality, multilingual corpus specifically designed for HSS emerging concepts, accompanied by a parallel English supervised dataset with fine-grained thematic labels. The resulting resource enables cross-lingual lexical variation analysis, training of multilingual NLP models, and empirical social science research. It is both reusable and extensible, effectively bridging a critical gap between domain-specific knowledge modeling and computational linguistics applications.

Build a multilingual corpus for studying emerging HSS conceptsCreate a machine-learning dataset from expert lexicon occurrencesEnable lexical analysis and NLP applications via reproducible resources

This work addresses the challenge of low-resource machine translation, where performance is hindered by the scarcity of high-quality parallel corpora. To this end, the authors propose LALITA, a novel framework that systematically leverages lexical and linguistic features of source sentences—such as syntactic complexity—to identify and select high-value training samples. Combined with synthetic data augmentation, this approach significantly reduces data requirements while improving translation quality. The method is evaluated across multiple low-resource languages, including Hindi, Odia, Nepali, Nynorsk, and German, consistently yielding performance gains across training sets ranging from 50K to 800K sentence pairs. Notably, LALITA achieves these improvements with over 50% less training data, demonstrating both its efficiency and strong generalization capability.

data curationlow-resource machine translationparallel corpus

This work addresses the challenge of efficiently and accurately detecting whether a given text is partially or fully contained—particularly via near verbatim copying—within massive web-scale corpora. To this end, the authors introduce FindMyText, an open-source tool that leverages document fingerprinting with a novel mechanism for identifying contiguous matching fingerprint sequences, explicitly capturing near-exact copied segments rather than relying on holistic text similarity. The system incorporates a distributed disk-based index to enable scalable processing. The study also establishes the first benchmark specifically designed for text containment tasks, demonstrating that FindMyText significantly outperforms existing methods across diverse datasets including arXiv, Wikipedia, and general web corpora, thereby validating its efficiency, robustness, and practical utility.

copyrighted material verificationdocument fingerprintingnear-verbatim detection

Hot Scholars

NH

Nizar Habash

Professor of Computer Science, New York University Abu Dhabi
Natural Language ProcessingComputational LinguisticsArtificial Intelligence
DI

David Ifeoluwa Adelani

McGill University and Mila - Quebec AI Institute and Canada CIFAR AI Chair
Natural language processingMultilingualityMultilingual NLPAfricaNLP
SH

Shamsuddeen Hassan Muhammad

Bayero University, Kano, & Google DeepMind Academic Fellow at Imperial College London
Natural Language ProcessingSentiment AnalysisAfricaNLPLow-resource NLP
IA

Idris Abdulmumin

Postdoctoral Fellow, DSFSI, University of Pretoria
Machine TranslationNeural Machine TranslationNatural Language ProcessingInternet Technology
BS

Benoît Sagot

Directeur de recherches at Inria, head of the ALMAnaCH team
NLPLanguage ModellingLow-resource LanguagesMachine Translation