How to scale your HEP ML models: A recipe for robust architecture comparisons at scale

๐Ÿ“… 2026-10-05
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the lack of systematic scaling law validation and architecture comparison in machine learning for high-energy physics by proposing a standardized pipeline for deriving robust scaling laws and evaluating design choices. Methodologically, it employs a multi-task Transformer architecture combined with scaling analysis techniques under both compute-constrained and data-constrained regimes, validated on large-scale ATLAS datasets. This work presents the first prediction of jointly optimal configurations for model size, training epochs, learning rate, and batch size under early stopping. Furthermore, it recovers an approximate square-root dependence between model and data sizes, demonstrates that auxiliary objectives effectively reduce the primary classification loss, and identifies the onset threshold of the power-law regime.
๐Ÿ“ Abstract
Much of the recent progress in machine learning domains such as language models has come from scaling laws that predict performance as a function of training effort. In high-energy physics (HEP) similar behavior has now been observed. To aid further study, we present a systematic procedure to derive robust scaling laws and compare design choices on the relevant budget axes for HEP tasks. We first validate the full scaling trajectory on toy problems and then apply the procedure to multi-task transformers on the ~11 billion-jet ATLAS JetSet2 dataset, in both the compute- and data-constrained regimes. For the latter, we predict, to the best of our knowledge for the first time, the jointly optimal model size, training horizon, learning rate and batch size under early stopping. At compute-optimal scaling, we recover a near-equal $\sqrt{C}$ dependence of model and dataset size, and find that auxiliary objectives lower the primary jet-classification loss at equal compute budget. Expanding the inputs toward lower-level data systematically lowers the loss while leaving the scaling exponent nearly unchanged. The onset of the power-law regime is itself set by scale: below a threshold in dataset size the loss carries little information about high-compute scaling, underscoring the value of large, high-quality full-simulation datasets as a foundation for scaling studies and the development of foundation models in HEP.
Problem

Research questions and friction points this paper is trying to address.

High-Energy Physics
Scaling Laws
Architecture Comparison
Compute-Optimal Scaling
Foundation Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Scaling Laws
High-Energy Physics
Multi-task Transformers
Compute-Optimal Scaling
Foundation Models
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
M
Matthias Vigl
Technical University of Munich
N
Nikita Pond
University College London
J
Jackson Barr
University College London
A
Alexander Froch
University of Geneva
D
Dan Guest
Humboldt-Universitรคt zu Berlin
N
Nicole Hartman
NSF AI Institute for Artificial Intelligence and Fundamental Interactions
M
Michael Kagan
SLAC National Accelerator Laboratory
Lukas Heinrich
Lukas Heinrich
Technical University of Munich
Particle PhysicsMachine LearningStatistics