🤖 AI Summary
This work addresses the substantial computational waste in hyperparameter optimization caused by training runs that are destined to fail early on. The authors propose a method to accurately predict final model performance using only telemetry data—such as loss, accuracy, gradient signal-to-noise ratio, weight norm dynamics, and activation saturation—from the first few epochs of a single training run, along with its hyperparameters, without requiring information from other runs. For the first time, they systematically validate the predictive power of early-training signals and demonstrate the incremental value of gradient- and weight-level metrics. Leveraging a gradient-boosted tree model across 23,788 experiments, they achieve R² values of 0.92–0.99 in predicting final accuracy and ROC-AUC scores of 0.983–0.998 for relative performance ranking using just the first five epochs, with useful predictive signals emerging as early as after a single epoch.
📝 Abstract
Large hyperparameter sweeps for deep neural networks spend substantial compute on configurations that are effectively doomed from the first few epochs. We study whether a single training run's own early telemetry - per-epoch loss, training accuracy, gradient signal-to-noise ratio, weight-norm growth, and an activation-saturation snapshot - together with its sampled hyperparameters, can predict that run's eventual outcome without reference to other runs. We evaluate three prediction tasks: final test accuracy, relative performance within a domain, and training-dynamics failure, including numerical divergence. Across 23,788 training runs spanning six architecture/dataset combinations, gradient-boosted trees using only the first five epochs of telemetry achieve R^2 = 0.92-0.99 for final-accuracy regression and ROC-AUC = 0.983-0.998 for relative classification on a permanently held-out set of hyperparameter configurations. Useful prediction is already available after a single epoch. A paired ablation shows that gradient- and weight-level telemetry provides a statistically consistent improvement over loss and accuracy curves alone, although the practical gain varies by domain. Transfer is strong between similar architectures, while cross-dataset transfer is limited mainly by differences in accuracy scale rather than loss of the underlying relationship. These results suggest that early-training telemetry can provide a practical decision-support signal for compute allocation while motivating human oversight for any automated intervention.