Small Experiments, Cheaper Decisions: A Case Study in Staged Promotion for Micro-Pretraining

📅 2026-06-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the risk of selecting suboptimal configurations under limited pretraining budgets, which can lead to significant resource waste. To mitigate this, the authors propose an auditable, staged promotion protocol that operates within a fixed micro-pretraining environment. By employing multi-stage time budgets ranging from 2 minutes to 12 hours, predefined promotion rules, replicated experiments across heterogeneous hardware (A100/L40S), and multiple random seeds, the method leverages Staged Factorial Screening and the val_bpb metric to distinguish genuine operational evidence of superiority from performance fluctuations due to single-seed variance. Strict near-equivalence and mean-gap criteria are applied to control cost and risk. The entire process consumes only 169.2 GPU-hours and successfully identifies a bridging configuration that consistently leads at the 12-hour stage, achieving over 60% GPU-hour savings compared to full-scale training.
📝 Abstract
Short pretraining runs can reduce experimental cost, but they can also over-promote configurations that only look strong at tiny budgets. We study an auditable staged-promotion protocol for a fixed micro-pretraining runner on two heterogeneous host blocks: Windows A100 and Linux L40S. Starting from twelve prior-screened configurations, we use staged budgets of 2 minutes, 5 minutes, 10 minutes, 60 minutes, and 12 hours, with frozen promotion rules before expensive continuations. The early screens are intentionally treated as unstable: the 5- and 10-minute rankings are host-sensitive, and the eventual 12-hour top-ranked condition is not the mean-best condition at the replicated 10-minute gate. Because seed ranges differ across stages, these changes are operational promotion evidence, not within-seed curves. A replicated 60-minute gate keeps the Staged Factorial Screening bridge reference in the promoted set, where it ranks first in all four 60-minute host-seed cells. In the final 12-hour confirmation package, the bridge condition ranks first in all four host-seed cells across two seeds; the greedy comparator does not meet the frozen 0.010 val_bpb near-equivalence rule; and the cheaper d8/ar48 (depth-8, aspect-48) sentinel does not meet the frozen 0.020 mean-gap rule. The executed 12-hour branch spends 144 GPU-hours, and the full staged protocol records 169.2 training GPU-hours including screening stages. Continuing all four 60-minute candidates would spend 192 GPU-hours, while continuing all nine replicated 10-minute candidates would spend 432 GPU-hours. The latter numbers are accounting counterfactuals for unrun continuations, not evidence that skipped candidates could not have overtaken the reference. The result is a bounded cost-allocation finding, not a claim of global optimality, capacity-normalized superiority, or superiority over adaptive hyperparameter optimization methods.
Problem

Research questions and friction points this paper is trying to address.

micro-pretraining
staged promotion
experimental cost
hyperparameter screening
budget allocation
Innovation

Methods, ideas, or system contributions that make the work stand out.

staged promotion
micro-pretraining
hyperparameter screening
GPU-hour efficiency
auditable protocol
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
F
Felipe Chavarro Polania
Hewlett Packard Enterprise