Scaling Laws for Classical Machine Learning on Tabular Data: A Benchmark Study

📅 2026-07-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of large-scale empirical analysis on scaling laws for classical machine learning models on tabular data. By coordinating 127 students to conduct 11,536 training runs across 18 datasets under a unified protocol, the authors fit power-law curves of the form error(N) = aN⁻ᵇ + c to characterize how prediction error decays with sample size N for six model families: Boosting, Random Forests, SVMs, linear/logistic regression, Ridge, and Lasso. This work presents the first multi-team, large-scale replication effort in this domain, introduces the concept of “approximate predictability compressibility,” and reveals that five out of six model families exhibit nearly shared exponents within each family. The study quantifies implementation-induced exponent variability (CV = 0.144), reports that 77.7% of fitted curves achieve R² > 0.8, and releases the complete dataset along with tables estimating the sample sizes required to reach target error levels.
📝 Abstract
Prior classical-ML learning-curve work fits power laws to tree, linear, and kernel models on tabular data, but at small scale: typically one curve, one team, a handful of cells. We present a distributed classroom-scale replication: 127 students each ran a fixed protocol on 3 assigned datasets, drawn from 18 tabular classification and regression datasets and 6 model families (Boosting, Random Forest, SVM, Linear/Logistic, Ridge, Lasso), yielding 11,536 training runs and 1,648 fitted power-law curves of the form error(N) = a N^(-b) + c. Three findings. (1) Power laws fit: R^2 > 0.8 on 77.7% of cells, with tree ensembles dominating at full data (Boosting 50% of datasets, RandomForest 33%; linear models underperform on classification). (2) Approximate shared exponents within a model family: for 5 of 6 families, a single family-level exponent predicts each family's cross-dataset curves nearly as well as per-dataset exponents (R^2 gap < 0.011), though AIC favors the unconstrained fit and curve collapse is partial (32-58% of points within +/-0.5 dex). We frame this as approximate predictive compressibility, not dataset-independent universality; Lasso fails outright (negative control) and Ridge is fragile under leave-one-dataset-out. (3) Replicator-implementation variance: with random_state=42 fixed, independent re-implementations of the same protocol still differ by mean CV(b) = 0.144 on the fitted exponent -- not seed variance, but the spread induced by unconstrained parts of the protocol (preprocessing, encoding, missing-value handling). We release the aggregated curves, per-cell fits, and a practical data-requirement table for N* to reach target error 0.15.
Problem

Research questions and friction points this paper is trying to address.

scaling laws
tabular data
classical machine learning
learning curves
power-law fitting
Innovation

Methods, ideas, or system contributions that make the work stand out.

scaling laws
tabular data
power-law learning curves
model family exponent
replication variance
💼 Related Jobs
No related jobs found.
K
Kaihua Ding
University of Pennsylvania