Adapting English Quality Classifiers for Multilingual LLM Pretraining Data Selection

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of high-quality data filtering for large language model (LLM) pretraining in low-resource languages, where annotated data is scarce. It proposes a cross-lingual quality classifier adaptation framework that eliminates the need for human annotations in target languages. By leveraging English quality scores of machine-translated texts as pseudo-labels and repurposing the MLP layers of an English classifier alongside Transformer encoder embeddings, the approach efficiently transfers the model to over one hundred languages. Experiments across 1B to 8B parameter scales demonstrate that this method effectively preserves benchmark performance without compromising region-specific cultural knowledge. Furthermore, it exhibits strong generalization to unseen languages, offering a universal and scalable solution for multilingual LLM data curation.
📝 Abstract
Recent advances in large language model (LLM) pretraining highlight the role of high-quality training data in improving performance. While model-based filtering has proven effective in selecting high-quality subsets from web-scale corpora, especially for high-resource languages, low-resource languages face challenges due to limited availability of annotated data. This work explores extending quality filtering to over 100 languages by proposing a multilingual adaptation approach that converts an existing English quality classifier into a multilingual variant. Our approach proposes training a small multi-layer perceptron on top of Transformer encoder-only model embeddings, using multilingual text as input and scores obtained from English classifiers applied to machine-translated text as labels. Our 1B, 3B and 8B scale experiments show that our approach maintains the downstream LLM benchmark performance of existing multilingual model-based filtering baselines, without harming regional and cultural knowledge benchmarks. To further evaluate cross-lingual generalization, we compare classifier scores of high-quality synthetic data and web samples, and the correlation of classifier scores with LLM-based ones, revealing that the classifier can learn the scoring criteria of its original English variant, even for languages not included in its training data.
Problem

Research questions and friction points this paper is trying to address.

Multilingual LLM
Data Selection
Quality Filtering
Low-resource Languages
Cross-lingual Generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multilingual Quality Filtering
Cross-lingual Transfer
Data Selection
Large Language Models
Machine Translation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.