Towards Scalable Data Diversification for Language Model Pretraining via Leverage Score Sampling

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the diversity collapse induced by data quality filtering in large model pre-training and the prohibitive computational costs of existing diversification methods. To this end, it proposes a scalable data selection approach based on leverage score sampling. By integrating determinantal point processes, the method iteratively selects samples that maximize the determinant volume in the embedding space, while employing leverage scores to circumvent expensive covariance matrix recomputation. This achieves, for the first time, large-scale selection of high-quality and diverse data. Experimental results demonstrate that the proposed method yields a 72× speedup over baselines, improves the Vendi score by 9.2%, increases downstream task accuracy by up to 1.31%, and reduces the compression ratio on code data by 3.08%.
📝 Abstract
Data selection for language model pretraining faces a fundamental tension between quality and diversity. While quality filtering is empirically effective, it often induces diversity collapse: by favoring texts similar to high-quality reference corpora (e.g., educational or QA-style data), it systematically excludes valuable data from underrepresented domains. In contrast, diversified selection preserves domain balance and encourages robust downstream performance, yet existing methods either focus on coverage-oriented objectives that indirectly enhance diversity, or directly optimize for diversity via costly covariance matrix recomputation that limits scalability. To address these issues, we introduce \textbf{Leverage Score Sampling (Lev)}, which iteratively selects samples that maximally expand the determinantal volume of the embedded data via leverage scores, a computationally efficient criterion that eliminates matrix recomputation and enables scalable selection. Empirically, Lev delivers up to $72\times$ speedup and improves dataset diversity, measured by the Vendi score, by $9.2\%$ over the strong diversification baseline \textbf{DiSF}. On CommonCrawl (CC) web data selection, Lev improves accuracy across seven downstream tasks by up to $1.31\%$ over existing baselines. For domains where robust quality criteria are inherently difficult to define (e.g., code), Lev serves as an effective unsupervised curation alternative: on StarCoderData, the selected subset reduces bits-per-byte by $3.08\%$ over DiSF. Notably, we uncover a cross-domain collapse of quality filtering: CC data filtered by DCLM-fastText fail to retain sufficient code-related content, yielding inferior code performance relative to Lev-selected data. These findings advocate for integrating diversity-aware practices into quality filtering for more effective data curation in language model pretraining.
Problem

Research questions and friction points this paper is trying to address.

data selection
language model pretraining
diversity collapse
scalability
quality-diversity tradeoff
Innovation

Methods, ideas, or system contributions that make the work stand out.

Leverage Score Sampling
Data Diversification
Language Model Pretraining
Determinantal Volume
Scalable Data Selection