When do data mixtures improve scaling laws? Insights from high-dimensional regression

๐Ÿ“… 2026-09-29
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses when data mixing genuinely improves scaling laws rather than merely increasing sample size. Leveraging high-dimensional regression theory, the authors analyze minimax risk under heterogeneous covariances and noise, deriving a deterministic equivalent for Ridge regression. They establish explicit scaling laws relating spectral decay, target regularity, and relative sample sizes to identify theoretical regimes where mixing outperforms single-domain training. The primary contribution lies in revealing sufficient conditions under which data mixing enhances scaling rates. Experiments on large language models confirm that, under specific conditions, appropriate data mixing enables target-domain test loss to decrease at a faster rate, significantly outperforming training on either dataset alone.
๐Ÿ“ Abstract
Modern machine learning systems are trained on mixtures of data from different domains, and choosing the right mixture can substantially improve downstream performance. Despite an extensive literature on data mixing and reweighting, existing work is largely empirical and it remains unclear when auxiliary data genuinely improves scaling laws rather than merely providing more samples. To gain insight into this question, we study a high-dimensional mixed-data regression model with a shared regression function, heterogeneous covariances and noise levels, and dataset sizes that may grow at different rates. We establish the minimax risk under an ellipsoidal parameter constraint for the general covariance structure and derive deterministic equivalents for the test error of ridge regression under commutative covariances. We then specialize to a target domain and an auxiliary domain with aligned power-law covariance spectra, where the theory yields explicit scaling laws in terms of spectral decay, target regularity, and the relative growth of the two datasets. These laws identify regimes in which combining data mixtures provably yields a faster scaling rate than using either dataset alone. In particular, improving the scaling law requires a specific interplay between spectra and relative sample sizes of the domains. Our numerical experiments on language models exhibit the same qualitative phenomenon: appropriate data mixtures yield a faster decrease in target-domain test loss than training on either domain alone.
Problem

Research questions and friction points this paper is trying to address.

data mixtures
scaling laws
high-dimensional regression
auxiliary data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Scaling Laws
High-dimensional Regression
Data Mixing
Ridge Regression
Minimax Risk