Score
Designs and implements statistical and algorithmic methods to fit scaling laws (e.g., power-law and multi-term/three-term forms) to empirical performance-versus-resource data, producing robust parameter estimates and uncertainty quantification even when some runs are suboptimal or limited. Uses these fitted laws for principled extrapolation and prediction across resource dimensions (such as batch size, model size, or compute) and for comparing or diagnosing experimental regimes.
Modern foundation models rely on scaling laws to guide training, yet existing studies exhibit inconsistent and irreproducible conclusions when extrapolating optimal architectures and hyperparameters—such as the tokens-to-parameters ratio—from small-scale experiments, due to heterogeneous fitting methodologies, divergent training configurations, and insufficient reporting of experimental details. Method: We systematically review over 50 scaling law papers and find that although 45 adopt power-law fitting, most omit critical experimental details; through controlled-variable empirical analysis, we demonstrate that minor configuration changes induce scaling exponent deviations exceeding 20%, substantially altering architectural recommendations. Contribution/Results: We propose the first standardized checklist for scaling law research and quantitatively characterize the sensitivity of fitting outcomes to multidimensional experimental variables, thereby establishing a methodological foundation to enhance the reliability and reproducibility of scaling studies.
Traditional neural scaling laws exhibit diminishing applicability in emerging architectures—including sparse models, Mixture-of-Experts (MoE), multimodal systems, and retrieval-augmented models—due to heterogeneity across modalities and stringent deployment constraints. Method: This work systematically reviews the theoretical foundations and empirical boundaries of scaling laws, synthesizing insights from over 50 studies. We propose an adaptive scaling framework that jointly optimizes data efficiency, inference cost, and architectural constraints, integrating power-law modeling, cross-modal performance attribution, and architecture-sensitivity analysis. Contribution/Results: We distill transferable, practical scaling guidelines that precisely delineate the validity domains and failure modes of scaling laws. The framework provides principled theoretical support and actionable decision-making tools for efficient large-model scaling; its components have been adopted by multiple industrial-scale training frameworks.
This work addresses the high cost of small-scale pilot experiments required to fit scaling laws in large-scale model training. Framing the problem as a budget-aware sequential experimental design task, the authors propose an uncertainty-aware active selection strategy that dynamically chooses the most informative experiments from a heterogeneous-cost pool for extrapolation to the target regime. By integrating sequential experimental design, uncertainty quantification, and active learning, the method achieves fitting accuracy comparable to that of exhaustive experimentation using only approximately 10% of the total training budget across diverse scaling law tasks, substantially outperforming conventional experimental design baselines.
Traditional scaling laws rely solely on model and data scale to predict performance, neglecting other hyperparameters and thus struggling to achieve accurate prediction and efficient tuning under hardware constraints. This work proposes Configuration-to-Performance Scaling Laws (CPL), which, for the first time, incorporate the full training configuration into the modeling framework. By parameterizing this mapping with a large language model, the authors introduce a neuralized CPL (NCPL). Trained on open-source pretraining logs, NCPL enables joint optimization across multiple hyperparameters and predicts loss curves with 20–40% lower error than Chinchilla scaling laws. It generalizes effectively to regimes up to ten times the maximum compute budget observed in the training set and matches baseline methods in multi-hyperparameter tuning tasks.
This work addresses the critical challenge of optimally allocating computational resources in large-scale model training by balancing the number of training steps against batch size. The authors propose a novel three-factor scaling law that explicitly decomposes the total training data into training steps and batch size, thereby unifying the modeling of how model scale, training steps, and batch size jointly influence performance. This formulation enables robust estimation of scaling relationships using data from suboptimal batch sizes and is validated through extensive large-scale experiments. The proposed scaling law accurately reproduces optimal-batch-size behavior, empirically confirms the existence of a critical batch size, and substantially reduces the number of required training trials, offering principled guidance for efficient resource allocation in practice.
This work addresses the scaling laws of hyperparameters—particularly learning rate, weight decay, and batch size—in large language model (LLM) pretraining with respect to model size (N), dataset size (D), and batch size (B). We propose a theoretical framework grounded in the time-scale invariance of AdamW: (B/(eta lambda D)). We establish that the optimal weight decay scales linearly with batch size ((lambda_{ ext{opt}} propto B)) and is precisely governed by a power law of the tokens-per-parameter ratio (D/N); furthermore, both the critical and optimal batch sizes scale as power laws of (D) alone, independent of (N). Through extensive controlled experiments and empirical validation of scaling exponents, we derive a Pareto-optimal joint selection strategy for (N) and (D) under dual objectives—training time and computational cost. The framework is validated across LLaMA and Pythia architectures, enabling high-accuracy *a priori* prediction of (lambda_{ ext{opt}}) and reducing hyperparameter tuning costs by orders of magnitude—e.g., millions of GPU-hours—thereby providing reusable, scalable configuration guidelines for thousand-GPU pretraining.
This work resolves a theoretical tension between neural scaling laws—which posit monotonic test error reduction with model size—and the classical bias-variance decomposition—which predicts increasing variance with parameter count. The authors analyze infinite-dimensional linear regression under a power-law spectral covariance structure and Gaussian prior, incorporating single-pass stochastic gradient descent (SGD) and its implicit regularization. For the first time, this framework theoretically reproduces neural scaling laws. They rigorously derive a reducible error bound of Θ(M^{−(a−1)} + N^{−(a−1)/a}), where M is model size and N is sample size, showing that implicit regularization dominates and suppresses the variance term—preventing its growth with model scale. Numerical experiments confirm the predicted power-law error decay. This provides a unified theoretical explanation for the persistent performance gains observed in increasingly large models.
This work addresses the common misconception that scaling laws apply only to large models, which often arises because small-scale models are evaluated with suboptimal hyperparameters. Through systematic analysis, the study demonstrates that scaling laws remain valid even for small models when evaluated along a properly tuned hyperparameter frontier, and further reveals that hyperparameter sensitivity diminishes as model scale increases. Building on these insights, the authors propose a new paradigm that combines small-scale experiments with efficient hyperparameter optimization. Using ablation studies, loss landscape analysis, and scaling modeling, they successfully reproduce findings typically observed only at large scales—such as the superiority of pre-normalization—thereby establishing that, under appropriate methodology, small-scale experiments can reliably predict large-model behavior.
Traditional scaling law estimation suffers from high computational costs due to the absence of efficient budget allocation strategies. This work proposes a novel approach that, for the first time, integrates surrogate-guided pruning into scaling law modeling by combining the Successive Halving algorithm with both parametric and non-parametric surrogate models. This integration enables proactive allocation of computational resources and efficient construction of loss-compute Pareto frontiers. The method substantially improves resource utilization efficiency, achieving relative performance gains of up to 2.84% on real datasets and 5.47% on synthetic datasets, while reducing computational costs by as much as 98.7%.
This study addresses the lack of systematic evaluation of statistical power among multivariate two-sample and goodness-of-fit tests, which hinders the selection of effective methods in practice. Through extensive Monte Carlo simulations, it presents the first comprehensive comparison of numerous nonparametric tests under bivariate settings—encompassing both continuous and discrete data—as well as high-dimensional continuous scenarios. Based on empirical findings, the paper proposes a small yet complementary ensemble of methods that collectively ensure high power against a wide range of alternative hypotheses. This ensemble demonstrates strong robustness and broad coverage, significantly outperforming any single test and offering practitioners a reliable, principled recommendation for real-world applications.
This study addresses the lack of large-scale empirical analysis on scaling laws for classical machine learning models on tabular data. By coordinating 127 students to conduct 11,536 training runs across 18 datasets under a unified protocol, the authors fit power-law curves of the form error(N) = aN⁻ᵇ + c to characterize how prediction error decays with sample size N for six model families: Boosting, Random Forests, SVMs, linear/logistic regression, Ridge, and Lasso. This work presents the first multi-team, large-scale replication effort in this domain, introduces the concept of “approximate predictability compressibility,” and reveals that five out of six model families exhibit nearly shared exponents within each family. The study quantifies implementation-induced exponent variability (CV = 0.144), reports that 77.7% of fitted curves achieve R² > 0.8, and releases the complete dataset along with tables estimating the sample sizes required to reach target error levels.