fit scaling laws robustly

Designs and implements statistical and algorithmic methods to fit scaling laws (e.g., power-law and multi-term/three-term forms) to empirical performance-versus-resource data, producing robust parameter estimates and uncertainty quantification even when some runs are suboptimal or limited. Uses these fitted laws for principled extrapolation and prediction across resource dimensions (such as batch size, model size, or compute) and for comparing or diagnosing experimental regimes.

fitscalinglawsrobustly

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

How to Upscale Neural Networks with Scaling Law? A Survey and Practical Guidelines

Feb 17, 2025
AS
Ayan Sengupta
🏛️ Indian Institute of Technology Delhi

Traditional neural scaling laws exhibit diminishing applicability in emerging architectures—including sparse models, Mixture-of-Experts (MoE), multimodal systems, and retrieval-augmented models—due to heterogeneity across modalities and stringent deployment constraints. Method: This work systematically reviews the theoretical foundations and empirical boundaries of scaling laws, synthesizing insights from over 50 studies. We propose an adaptive scaling framework that jointly optimizes data efficiency, inference cost, and architectural constraints, integrating power-law modeling, cross-modal performance attribution, and architecture-sensitivity analysis. Contribution/Results: We distill transferable, practical scaling guidelines that precisely delineate the validity domains and failure modes of scaling laws. The framework provides principled theoretical support and actionable decision-making tools for efficient large-model scaling; its components have been adopted by multiple industrial-scale training frameworks.

Adapting scaling laws across architecturesAddressing domain-specific scaling challengesUpscaling neural networks efficiently

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the high cost of small-scale pilot experiments required to fit scaling laws in large-scale model training. Framing the problem as a budget-aware sequential experimental design task, the authors propose an uncertainty-aware active selection strategy that dynamically chooses the most informative experiments from a heterogeneous-cost pool for extrapolation to the target regime. By integrating sequential experimental design, uncertainty quantification, and active learning, the method achieves fitting accuracy comparable to that of exhaustive experimentation using only approximately 10% of the total training budget across diverse scaling law tasks, substantially outperforming conventional experimental design baselines.

active learningbudget efficiencyexperimental design

Traditional scaling laws rely solely on model and data scale to predict performance, neglecting other hyperparameters and thus struggling to achieve accurate prediction and efficient tuning under hardware constraints. This work proposes Configuration-to-Performance Scaling Laws (CPL), which, for the first time, incorporate the full training configuration into the modeling framework. By parameterizing this mapping with a large language model, the authors introduce a neuralized CPL (NCPL). Trained on open-source pretraining logs, NCPL enables joint optimization across multiple hyperparameters and predicts loss curves with 20–40% lower error than Chinchilla scaling laws. It generalizes effectively to regimes up to ten times the maximum compute budget observed in the training set and matches baseline methods in multi-hyperparameter tuning tasks.

hyperparameter tuninglarge language modelsperformance prediction

This work addresses the critical challenge of optimally allocating computational resources in large-scale model training by balancing the number of training steps against batch size. The authors propose a novel three-factor scaling law that explicitly decomposes the total training data into training steps and batch size, thereby unifying the modeling of how model scale, training steps, and batch size jointly influence performance. This formulation enables robust estimation of scaling relationships using data from suboptimal batch sizes and is validated through extensive large-scale experiments. The proposed scaling law accurately reproduces optimal-batch-size behavior, empirically confirms the existence of a critical batch size, and substantially reduces the number of required training trials, offering principled guidance for efficient resource allocation in practice.

batch sizemodel scalingoptimal allocation

This work addresses the scaling laws of hyperparameters—particularly learning rate, weight decay, and batch size—in large language model (LLM) pretraining with respect to model size (N), dataset size (D), and batch size (B). We propose a theoretical framework grounded in the time-scale invariance of AdamW: (B/(eta lambda D)). We establish that the optimal weight decay scales linearly with batch size ((lambda_{ ext{opt}} propto B)) and is precisely governed by a power law of the tokens-per-parameter ratio (D/N); furthermore, both the critical and optimal batch sizes scale as power laws of (D) alone, independent of (N). Through extensive controlled experiments and empirical validation of scaling exponents, we derive a Pareto-optimal joint selection strategy for (N) and (D) under dual objectives—training time and computational cost. The framework is validated across LLaMA and Pythia architectures, enabling high-accuracy *a priori* prediction of (lambda_{ ext{opt}}) and reducing hyperparameter tuning costs by orders of magnitude—e.g., millions of GPU-hours—thereby providing reusable, scalable configuration guidelines for thousand-GPU pretraining.

Analyze Pareto-optimal model and dataset size selectionDetermine optimal weight decay and batch size scalingStudy scaling laws for hyperparameters in LLM pre-training

Scaling Laws in Linear Regression: Compute, Parameters, and Data

Jun 12, 2024
LL
Licong Lin
🏛️ UC Berkeley | Harvard University | Google DeepMind | Princeton University

This work resolves a theoretical tension between neural scaling laws—which posit monotonic test error reduction with model size—and the classical bias-variance decomposition—which predicts increasing variance with parameter count. The authors analyze infinite-dimensional linear regression under a power-law spectral covariance structure and Gaussian prior, incorporating single-pass stochastic gradient descent (SGD) and its implicit regularization. For the first time, this framework theoretically reproduces neural scaling laws. They rigorously derive a reducible error bound of Θ(M^{−(a−1)} + N^{−(a−1)/a}), where M is model size and N is sample size, showing that implicit regularization dominates and suppresses the variance term—preventing its growth with model scale. Numerical experiments confirm the predicted power-law error decay. This provides a unified theoretical explanation for the persistent performance gains observed in increasingly large models.

Analyze test error reducible part componentsResolve variance error increase with model sizeUnderstand scaling laws in linear regression models

Latest Papers

What's happening recently
View more

This work addresses the common misconception that scaling laws apply only to large models, which often arises because small-scale models are evaluated with suboptimal hyperparameters. Through systematic analysis, the study demonstrates that scaling laws remain valid even for small models when evaluated along a properly tuned hyperparameter frontier, and further reveals that hyperparameter sensitivity diminishes as model scale increases. Building on these insights, the authors propose a new paradigm that combines small-scale experiments with efficient hyperparameter optimization. Using ablation studies, loss landscape analysis, and scaling modeling, they successfully reproduce findings typically observed only at large scales—such as the superiority of pre-normalization—thereby establishing that, under appropriate methodology, small-scale experiments can reliably predict large-model behavior.

hyperparameter sensitivitymodel scalingscaling laws

Traditional scaling law estimation suffers from high computational costs due to the absence of efficient budget allocation strategies. This work proposes a novel approach that, for the first time, integrates surrogate-guided pruning into scaling law modeling by combining the Successive Halving algorithm with both parametric and non-parametric surrogate models. This integration enables proactive allocation of computational resources and efficient construction of loss-compute Pareto frontiers. The method substantially improves resource utilization efficiency, achieving relative performance gains of up to 2.84% on real datasets and 5.47% on synthetic datasets, while reducing computational costs by as much as 98.7%.

compute budget allocationefficient estimationlearning curves

This study addresses the lack of systematic evaluation of statistical power among multivariate two-sample and goodness-of-fit tests, which hinders the selection of effective methods in practice. Through extensive Monte Carlo simulations, it presents the first comprehensive comparison of numerous nonparametric tests under bivariate settings—encompassing both continuous and discrete data—as well as high-dimensional continuous scenarios. Based on empirical findings, the paper proposes a small yet complementary ensemble of methods that collectively ensure high power against a wide range of alternative hypotheses. This ensemble demonstrates strong robustness and broad coverage, significantly outperforming any single test and offering practitioners a reliable, principled recommendation for real-world applications.

goodness-of-fitmultivariate datanon-parametric

This study addresses the lack of large-scale empirical analysis on scaling laws for classical machine learning models on tabular data. By coordinating 127 students to conduct 11,536 training runs across 18 datasets under a unified protocol, the authors fit power-law curves of the form error(N) = aN⁻ᵇ + c to characterize how prediction error decays with sample size N for six model families: Boosting, Random Forests, SVMs, linear/logistic regression, Ridge, and Lasso. This work presents the first multi-team, large-scale replication effort in this domain, introduces the concept of “approximate predictability compressibility,” and reveals that five out of six model families exhibit nearly shared exponents within each family. The study quantifies implementation-induced exponent variability (CV = 0.144), reports that 77.7% of fitted curves achieve R² > 0.8, and releases the complete dataset along with tables estimating the sample sizes required to reach target error levels.

classical machine learninglearning curvespower-law fitting

Hot Scholars

SQ

Shikai Qiu

PhD Student, New York University
Deep Learning
CT

Changxin Tian

Renmin University of China & Ant Group
Large Language Models
NL

Noam Levi

Postdoctoral Fellow, AI4Science/AI Center, EPFL
RMTStatistical Learning TheoryField TheoryParticle Physics
PP

Paolo Paradisi

Research scientist, ISTI-CNR,Pisa
complexitystatistical physicsneural networks and neurosciencelearning
BN

Benjamin Nachman

Staff Scientist, Lawrence Berkeley National Laboratory
Particle PhysicsDeep LearningQuantum ComputingSolid State Detectors