Score
Designs, builds, and curates text and multimodal corpora — including bitext/parallel and multi‑parallel corpora, synthetic or mass‑translated pretraining datasets, and aligned image–text collections — by collecting, mining, sampling, aligning (e.g., sentence/row alignment), annotating, and preprocessing raw data for training or evaluation. Evaluates and validates corpus quality and coverage through filtering, denoising, normalization, language identification, register balancing, and other corpus‑linguistic analyses to ensure fitness for downstream modeling or linguistic study.
This study systematically investigates optimal strategies for leveraging parallel corpora to enhance multilingual large language models (MLLMs), focusing on the impacts of corpus quality and scale, training objectives, and model parameter count on both bilingual tasks (e.g., machine translation) and general cross-lingual tasks (e.g., text classification). Method: We propose a noise-filtering–based parallel corpus selection mechanism—bypassing error-prone language identification preprocessing—and employ supervised fine-tuning with a pure machine translation (MT) objective, integrated with multilingual pretraining. Contribution/Results: We find that merely ~10K high-quality parallel sentence pairs achieve performance comparable to large-scale corpora; the MT-only objective substantially outperforms multitask mixed objectives; and larger models benefit more markedly from parallel data. Experiments demonstrate consistent improvements across 12 languages and five cross-lingual task categories, establishing a reusable, efficient paradigm for parallel corpus utilization in MLLM development.
This study addresses the lack of systematic approaches for constructing, storing, and sharing high-quality annotated corpora. It proposes a generalizable and reusable end-to-end methodology encompassing annotation guideline development, corpus annotation, data storage, sharing mechanisms, and value realization, with an emphasis on full lifecycle management and cross-domain applicability. Integrating linguistic annotation theory, data management standards, and collaborative research practices, the approach is articulated through a structured framework and illustrative examples to yield a clear and actionable guide. The resulting methodology provides standardized support for diverse research domains, significantly enhancing the efficiency and quality with which researchers can build and utilize annotated textual data.
This work addresses the challenge of low-resource machine translation, where performance is hindered by the scarcity of high-quality parallel corpora. To this end, the authors propose LALITA, a novel framework that systematically leverages lexical and linguistic features of source sentences—such as syntactic complexity—to identify and select high-value training samples. Combined with synthetic data augmentation, this approach significantly reduces data requirements while improving translation quality. The method is evaluated across multiple low-resource languages, including Hindi, Odia, Nepali, Nynorsk, and German, consistently yielding performance gains across training sets ranging from 50K to 800K sentence pairs. Notably, LALITA achieves these improvements with over 50% less training data, demonstrating both its efficiency and strong generalization capability.
In text classification, manual verification of predictions is costly and ill-suited for continuous retraining under data drift. This work pioneers a systematic investigation into leveraging large language models (LLMs) as trustworthy automated validators—replacing human annotation to ensure classifier quality and enable efficient incremental updates. Our method integrates prompt engineering, zero- and few-shot inference, consistency checking, task-specific semantic constraints, and model confidence analysis. Evaluated across multiple benchmark datasets, LLM-based validation achieves over 92% agreement with expert annotations, substantially reducing verification cost while improving pipeline timeliness and scalability. The core contribution is the first LLM-based trustworthy validation framework specifically designed for classifier prediction verification—establishing a new paradigm for low-cost, robust continual learning.
To address the scarcity of pretraining data for low-resource Indian languages, this paper introduces BhashaKritika, a multilingual synthetic data construction framework covering 10 Indian languages and 54 billion tokens. Methodologically, it proposes the first document–role–topic co-guided synthetic generation paradigm, integrating five complementary generation techniques; it further designs a modular quality assurance pipeline incorporating script/language identification, metadata consistency verification, n-gram deduplication, and KenLM perplexity filtering—enabling efficient cross-script and cross-lingual quality control. Comprehensive experiments systematically characterize the quality–diversity trade-offs across generation strategies, establishing best practices for multilingual synthetic corpus construction. Empirical results demonstrate that models pretrained on BhashaKritika achieve substantial performance gains across Indian languages, providing a reusable data infrastructure and methodological blueprint for low-resource multilingual LLM development.
This study addresses the critical challenges in Lombard, a low-resource language, where existing NLP corpora suffer from widespread mislabeling, templated content, noise, and severe dialectal imbalance. For the first time, the authors conduct a systematic manual audit, integrating language identification validation, orthographic analysis, and dialect classification to assess linguistic authenticity and regional representativeness. Their findings reveal that mainstream datasets contain an extremely low proportion of genuine Lombard texts, exhibit inconsistent orthography, and display a pronounced representational bias favoring Western dialects while marginalizing Eastern varieties. The work underscores the necessity of moving beyond quantity-driven data collection toward community-informed, dialect-sensitive strategies for building high-quality linguistic resources, offering crucial methodological guidance for corpus development in other low-resource languages.
This work addresses the scarcity of large-scale, high-quality sentence-aligned corpora for text simplification in non-English languages by systematically constructing and publicly releasing a multilingual simplification corpus covering Catalan, English, French, Italian, and Spanish. Leveraging crowdsourcing, the authors collect simplified texts from comparable documents and implement a document-to-sentence alignment mechanism to produce high-quality sentence pairs. This resource fills a critical gap in non-English simplification data and provides a foundational benchmark for training and evaluating multilingual text simplification systems.
This work addresses the limited cross-lingual alignment in existing multilingual pretrained models, which stems from the absence of explicit alignment signals. To overcome this, the authors systematically leverage multidirectional parallel corpora—covering six languages and generated using off-the-shelf neural machine translation models—and fine-tune models such as XLM-R, mBERT, and mE5 via contrastive learning. This approach substantially enhances cross-lingual representation quality. Compared to conventional English-centric bilingual data, it yields significant improvements on the MTEB benchmark across text retrieval (+21.3%), semantic similarity (+5.3%), and classification tasks (+28.4%). Notably, the gains extend to unseen languages, demonstrating the unique efficacy of multidirectional parallel data for cross-lingual alignment.
This study addresses the scarcity of multilingual corpora supporting emerging concepts—such as “non-technological innovation”—in the humanities and social sciences (HSS). To tackle this, we propose a hybrid multilingual corpus construction methodology that integrates corporate websites and annual reports, combining automatic language identification, domain-adapted content filtering, relevant paragraph extraction, expert-lexicon-driven contextual block identification, thematic annotation, and enriched structured metadata. Our key contribution is the first systematic construction of a high-quality, multilingual corpus specifically designed for HSS emerging concepts, accompanied by a parallel English supervised dataset with fine-grained thematic labels. The resulting resource enables cross-lingual lexical variation analysis, training of multilingual NLP models, and empirical social science research. It is both reusable and extensible, effectively bridging a critical gap between domain-specific knowledge modeling and computational linguistics applications.
This study investigates the role of bilingual documents in the pretraining of multilingual large language models and their contribution to cross-lingual capabilities. By training models from scratch under controlled conditions—comparing standard corpora against strictly monolingual corpora with all multilingual documents removed—the authors quantitatively demonstrate for the first time that bilingual documents, though constituting only 2% of the training data, are critical for translation performance: their removal causes a 56% drop in BLEU score, while other cross-lingual tasks remain largely unaffected. Further analysis reveals that parallel texts, through token-level alignment, recover 91% of translation performance, whereas code-switched texts contribute minimally. These findings indicate that high-quality translation relies heavily on explicit alignment signals, whereas cross-lingual understanding can emerge without direct exposure to bilingual data.