text corpus construction

Designs, builds, and curates text and multimodal corpora — including bitext/parallel and multi‑parallel corpora, synthetic or mass‑translated pretraining datasets, and aligned image–text collections — by collecting, mining, sampling, aligning (e.g., sentence/row alignment), annotating, and preprocessing raw data for training or evaluation. Evaluates and validates corpus quality and coverage through filtering, denoising, normalization, language identification, register balancing, and other corpus‑linguistic analyses to ensure fitness for downstream modeling or linguistic study.

textcorpusconstruction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.4
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$222K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A Recipe of Parallel Corpora Exploitation for Multilingual Large Language Models

Jun 29, 2024
PL
Peiqin Lin
🏛️ LMU Munich | Instituto Superior Técnico | Universidade de Lisboa | Instituto de Telecomunicações | Unbabel

This study systematically investigates optimal strategies for leveraging parallel corpora to enhance multilingual large language models (MLLMs), focusing on the impacts of corpus quality and scale, training objectives, and model parameter count on both bilingual tasks (e.g., machine translation) and general cross-lingual tasks (e.g., text classification). Method: We propose a noise-filtering–based parallel corpus selection mechanism—bypassing error-prone language identification preprocessing—and employ supervised fine-tuning with a pure machine translation (MT) objective, integrated with multilingual pretraining. Contribution/Results: We find that merely ~10K high-quality parallel sentence pairs achieve performance comparable to large-scale corpora; the MT-only objective substantially outperforms multitask mixed objectives; and larger models benefit more markedly from parallel data. Experiments demonstrate consistent improvements across 12 languages and five cross-lingual task categories, establishing a reusable, efficient paradigm for parallel corpus utilization in MLLM development.

Enhancing large language models with parallel dataImproving performance across diverse tasksOptimizing parallel corpora for multilingual models

This study addresses the lack of systematic approaches for constructing, storing, and sharing high-quality annotated corpora. It proposes a generalizable and reusable end-to-end methodology encompassing annotation guideline development, corpus annotation, data storage, sharing mechanisms, and value realization, with an emphasis on full lifecycle management and cross-domain applicability. Integrating linguistic annotation theory, data management standards, and collaborative research practices, the approach is articulated through a structured framework and illustrative examples to yield a clear and actionable guide. The resulting methodology provides standardized support for diverse research domains, significantly enhancing the efficiency and quality with which researchers can build and utilize annotated textual data.

annotated corpusannotation guidelinescorpus creation

This work addresses the challenge of low-resource machine translation, where performance is hindered by the scarcity of high-quality parallel corpora. To this end, the authors propose LALITA, a novel framework that systematically leverages lexical and linguistic features of source sentences—such as syntactic complexity—to identify and select high-value training samples. Combined with synthetic data augmentation, this approach significantly reduces data requirements while improving translation quality. The method is evaluated across multiple low-resource languages, including Hindi, Odia, Nepali, Nynorsk, and German, consistently yielding performance gains across training sets ranging from 50K to 800K sentence pairs. Notably, LALITA achieves these improvements with over 50% less training data, demonstrating both its efficiency and strong generalization capability.

data curationlow-resource machine translationparallel corpus

In text classification, manual verification of predictions is costly and ill-suited for continuous retraining under data drift. This work pioneers a systematic investigation into leveraging large language models (LLMs) as trustworthy automated validators—replacing human annotation to ensure classifier quality and enable efficient incremental updates. Our method integrates prompt engineering, zero- and few-shot inference, consistency checking, task-specific semantic constraints, and model confidence analysis. Evaluated across multiple benchmark datasets, LLM-based validation achieves over 92% agreement with expert annotations, substantially reducing verification cost while improving pipeline timeliness and scalability. The core contribution is the first LLM-based trustworthy validation framework specifically designed for classifier prediction verification—establishing a new paradigm for low-cost, robust continual learning.

Addressing high costs and limited availability of human annotatorsAutomating text classifier validation using LLMs to reduce human effortEnsuring model quality and enabling efficient incremental learning

To address the scarcity of pretraining data for low-resource Indian languages, this paper introduces BhashaKritika, a multilingual synthetic data construction framework covering 10 Indian languages and 54 billion tokens. Methodologically, it proposes the first document–role–topic co-guided synthetic generation paradigm, integrating five complementary generation techniques; it further designs a modular quality assurance pipeline incorporating script/language identification, metadata consistency verification, n-gram deduplication, and KenLM perplexity filtering—enabling efficient cross-script and cross-lingual quality control. Comprehensive experiments systematically characterize the quality–diversity trade-offs across generation strategies, establishing best practices for multilingual synthetic corpus construction. Empirical results demonstrate that models pretrained on BhashaKritika achieve substantial performance gains across Indian languages, providing a reusable data infrastructure and methodological blueprint for low-resource multilingual LLM development.

Comparing translation-based and native generation approaches for Indic language corporaEvaluating multilingual data quality across diverse scripts and linguistic contextsGenerating scalable synthetic pretraining data for low-resource Indic languages

Latest Papers

What's happening recently
View more

This study addresses the critical challenges in Lombard, a low-resource language, where existing NLP corpora suffer from widespread mislabeling, templated content, noise, and severe dialectal imbalance. For the first time, the authors conduct a systematic manual audit, integrating language identification validation, orthographic analysis, and dialect classification to assess linguistic authenticity and regional representativeness. Their findings reveal that mainstream datasets contain an extremely low proportion of genuine Lombard texts, exhibit inconsistent orthography, and display a pronounced representational bias favoring Western dialects while marginalizing Eastern varieties. The work underscores the necessity of moving beyond quantity-driven data collection toward community-informed, dialect-sensitive strategies for building high-quality linguistic resources, offering crucial methodological guidance for corpus development in other low-resource languages.

corpus auditinglanguage identificationLombard

This work addresses the scarcity of large-scale, high-quality sentence-aligned corpora for text simplification in non-English languages by systematically constructing and publicly releasing a multilingual simplification corpus covering Catalan, English, French, Italian, and Spanish. Leveraging crowdsourcing, the authors collect simplified texts from comparable documents and implement a document-to-sentence alignment mechanism to produce high-quality sentence pairs. This resource fills a critical gap in non-English simplification data and provides a foundational benchmark for training and evaluating multilingual text simplification systems.

dataset scarcitymultilingual corporanon-English languages

This work addresses the limited cross-lingual alignment in existing multilingual pretrained models, which stems from the absence of explicit alignment signals. To overcome this, the authors systematically leverage multidirectional parallel corpora—covering six languages and generated using off-the-shelf neural machine translation models—and fine-tune models such as XLM-R, mBERT, and mE5 via contrastive learning. This approach substantially enhances cross-lingual representation quality. Compared to conventional English-centric bilingual data, it yields significant improvements on the MTEB benchmark across text retrieval (+21.3%), semantic similarity (+5.3%), and classification tasks (+28.4%). Notably, the gains extend to unseen languages, demonstrating the unique efficacy of multidirectional parallel data for cross-lingual alignment.

cross-lingual alignmentmulti-way parallel corpusmultilingual embeddings

Multilingual corpora for the study of new concepts in the social sciences and humanities:

Dec 08, 2025
RK
Revekka Kyriakoglou
🏛️ LIASD | Université Paris 8

This study addresses the scarcity of multilingual corpora supporting emerging concepts—such as “non-technological innovation”—in the humanities and social sciences (HSS). To tackle this, we propose a hybrid multilingual corpus construction methodology that integrates corporate websites and annual reports, combining automatic language identification, domain-adapted content filtering, relevant paragraph extraction, expert-lexicon-driven contextual block identification, thematic annotation, and enriched structured metadata. Our key contribution is the first systematic construction of a high-quality, multilingual corpus specifically designed for HSS emerging concepts, accompanied by a parallel English supervised dataset with fine-grained thematic labels. The resulting resource enables cross-lingual lexical variation analysis, training of multilingual NLP models, and empirical social science research. It is both reusable and extensible, effectively bridging a critical gap between domain-specific knowledge modeling and computational linguistics applications.

Build a multilingual corpus for studying emerging HSS conceptsCreate a machine-learning dataset from expert lexicon occurrencesEnable lexical analysis and NLP applications via reproducible resources

This study investigates the role of bilingual documents in the pretraining of multilingual large language models and their contribution to cross-lingual capabilities. By training models from scratch under controlled conditions—comparing standard corpora against strictly monolingual corpora with all multilingual documents removed—the authors quantitatively demonstrate for the first time that bilingual documents, though constituting only 2% of the training data, are critical for translation performance: their removal causes a 56% drop in BLEU score, while other cross-lingual tasks remain largely unaffected. Further analysis reveals that parallel texts, through token-level alignment, recover 91% of translation performance, whereas code-switched texts contribute minimally. These findings indicate that high-quality translation relies heavily on explicit alignment signals, whereas cross-lingual understanding can emerge without direct exposure to bilingual data.

bilingual datacode-switchingcross-lingual performance

Hot Scholars

BP

Barbara Plank

Professor, LMU Munich, Visiting Prof ITU Copenhagen
Natural Language ProcessingComputational LinguisticsMachine LearningTransfer Learning
EL

Eric Laporte

Université Gustave Eiffel
Linguistic description for language processingInformation retrieval
AZ

Amir Zeldes

Associate Professor of Computational Linguistics, Georgetown University
corpus linguisticscomputational linguisticsNLPdiscourse
SH

Shamsuddeen Hassan Muhammad

Bayero University, Kano, & Google DeepMind Academic Fellow at Imperial College London
Natural Language ProcessingSentiment AnalysisAfricaNLPLow-resource NLP
TS

Thales Sales Almeida

Student, Unicamp
Information retrievalMachine learningDeep learningGenerative models