Predicting Deep Neural Network Training Outcomes from Early Training Telemetry

📅 2026-08-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the substantial computational waste in hyperparameter optimization caused by training runs that are destined to fail early on. The authors propose a method to accurately predict final model performance using only telemetry data—such as loss, accuracy, gradient signal-to-noise ratio, weight norm dynamics, and activation saturation—from the first few epochs of a single training run, along with its hyperparameters, without requiring information from other runs. For the first time, they systematically validate the predictive power of early-training signals and demonstrate the incremental value of gradient- and weight-level metrics. Leveraging a gradient-boosted tree model across 23,788 experiments, they achieve R² values of 0.92–0.99 in predicting final accuracy and ROC-AUC scores of 0.983–0.998 for relative performance ranking using just the first five epochs, with useful predictive signals emerging as early as after a single epoch.
📝 Abstract
Large hyperparameter sweeps for deep neural networks spend substantial compute on configurations that are effectively doomed from the first few epochs. We study whether a single training run's own early telemetry - per-epoch loss, training accuracy, gradient signal-to-noise ratio, weight-norm growth, and an activation-saturation snapshot - together with its sampled hyperparameters, can predict that run's eventual outcome without reference to other runs. We evaluate three prediction tasks: final test accuracy, relative performance within a domain, and training-dynamics failure, including numerical divergence. Across 23,788 training runs spanning six architecture/dataset combinations, gradient-boosted trees using only the first five epochs of telemetry achieve R^2 = 0.92-0.99 for final-accuracy regression and ROC-AUC = 0.983-0.998 for relative classification on a permanently held-out set of hyperparameter configurations. Useful prediction is already available after a single epoch. A paired ablation shows that gradient- and weight-level telemetry provides a statistically consistent improvement over loss and accuracy curves alone, although the practical gain varies by domain. Transfer is strong between similar architectures, while cross-dataset transfer is limited mainly by differences in accuracy scale rather than loss of the underlying relationship. These results suggest that early-training telemetry can provide a practical decision-support signal for compute allocation while motivating human oversight for any automated intervention.
Problem

Research questions and friction points this paper is trying to address.

hyperparameter optimization
training outcome prediction
early stopping
deep neural networks
telemetry
Innovation

Methods, ideas, or system contributions that make the work stand out.

early-training telemetry
hyperparameter prediction
training dynamics
gradient signal-to-noise ratio
compute allocation
🔎 Similar Papers
No similar papers found.
R
Ranjita Naik
Georgia Institute of Technology
A
Anh D. Nguyen
Georgia Institute of Technology
P
Pankaj Kumar Singh
Georgia Institute of Technology