adapt training to architectures

Designs and implements training procedures, hyperparameter schedules, and optimizer configurations tailored to particular neural network architectures. This includes modifying loss functions, reparameterizing modules to preserve end-to-end gradient flow, and validating training outcomes and performance across different architectures.

adapttrainingtoarchitectures

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This paper addresses the inefficiency and lack of scalability of manual hyperparameter tuning in large-scale machine learning. It systematically surveys hyperparameter optimization (HPO), unifying and classifying five mainstream paradigms: random/low-discrepancy search, bandit-based methods, Bayesian optimization, population-based (evolutionary) algorithms, and gradient-based differentiable optimization. The survey further extends to emerging settings—including online HPO, constrained HPO, and multi-objective HPO. Crucially, the work establishes novel theoretical connections between HPO and meta-learning as well as neural architecture search, yielding a comprehensive knowledge framework that articulates methodological principles, applicability boundaries, and inherent limitations. By clarifying the technical evolution and identifying key open challenges, this study provides a theoretically grounded yet practically actionable foundation for automated machine learning.

Addressing challenges in online, constrained, and multi-objective hyperparameter tuningAutomating hyperparameter search to improve machine learning efficiencyComparing state-of-the-art hyperparameter optimization techniques and methods

Training Neural Networks at Any Scale

Nov 14, 2025
TP
Thomas Pethick
🏛️ École Polytechnique Fédérale de Lausanne | CVN | CentraleSupélec | Université Paris-Saclay | Inria | Gatsby Computational Neuroscience Unit | University College London

To address the low training efficiency, poor generalization, and strong hyperparameter sensitivity of neural networks across varying scales, this paper proposes a scale-invariant adaptive optimization framework. The method unifies adaptive optimization, second-order information approximation, learning-rate scaling invariance, and gradient compression, thereby decoupling optimization from model size and hardware configuration. Its core innovation lies in a scale-robust update paradigm that ensures stable optimization dynamics under variations in parameter count, batch size, and device count. Extensive experiments across diverse architectures—including MLPs, CNNs, and Transformers—and benchmarks—including CIFAR-10/100, ImageNet, and WikiText—demonstrate that the framework achieves 1.3–2.1× speedup over baseline optimizers, improved convergence stability, significantly reduced hyperparameter sensitivity, and eliminates the need for scale-specific hyperparameter tuning.

Developing algorithms adaptable to problem structuresMaking optimization methods independent of problem sizeOptimizing neural network training for efficiency and scale

This paper addresses the challenge of hyperparameter transferability across model scales—specifically width, depth, batch size, and training duration—where existing methods fail to generalize reliably. We propose Complete<sup>(d)</sup>, the first modular hyperparameter parameterization framework enabling robust cross-scale transfer under *multi-dimensional* coordinated scaling, overcoming the limitation of prior μP approaches that support only single-dimension scaling. Our method integrates modular hyperparameter optimization, AdamW hyperparameter space modeling, and joint search of residual block multipliers and initialization scales. Evaluated on large language models, Complete<sup>(d)</sup> demonstrates full-stack hyperparameter transferability—including learning rate, AdamW parameters (β₁, β₂, ε), weight decay, initialization scale, and residual scaling—across diverse model sizes. This yields substantially improved training stability and faster convergence, establishing a systematic, scalable paradigm for hyperparameter transfer in large-model training.

Demonstrates training speed improvements in Large Language ModelsExtends hyperparameter transfer to width, depth, batch, and duration scalingInvestigates per-module hyperparameter optimization and transfer challenges

Accelerating Neural Network Training: An Analysis of the AlgoPerf Competition

Feb 20, 2025
PK
Priya Kasimbeg
🏛️ Google DeepMind | University of Tübingen | University of Cambridge | Vector Institute | University of Toronto | Dalhousie University | Meta

This study addresses the fundamental question: “Can purely algorithmic improvements yield practical acceleration in neural network training?” To this end, we organized the inaugural AlgoPerf competition, establishing— for the first time—two rigorous evaluation paradigms: workload-agnostic assessment and hyperparameter-free benchmarking, with end-to-end training time on identical hardware as the sole primary metric. Methodologically, we developed a multi-task benchmarking framework integrating Distributed Shampoo (a non-diagonal preconditioner) and Schedule-Free AdamW (a hyperparameter-free optimizer), complemented by standardized temporal measurement protocols and fairness-preserving engineering safeguards. Results show that Distributed Shampoo achieved top performance in the hyperparameter-tuned track, while Schedule-Free AdamW led in the hyperparameter-free track. Top-performing methods demonstrated consistent speedups across diverse CV and NLP tasks, empirically validating that high-quality algorithmic design delivers substantial and robust training acceleration.

Hyperparameter-free training algorithmsNeural network training speed-upRobustness to workload changes

Benchmarking Neural Network Training Algorithms

Jun 12, 2023
GE
George E. Dahl
🏛️ Google | University of Tübingen | Vector Institute | Dalhousie University | University of Toronto | Stanford University | Meta AI | Dell Technologies

Fair evaluation of deep learning training algorithms faces three key challenges: inconsistent termination criteria, high workload sensitivity, and difficulty isolating hyperparameter tuning. This paper introduces AlgoPerf—the first time-oriented, multi-workload training algorithm benchmark—featuring robustness-aware workload variant design and a standardized termination protocol, with hyperparameter tuning rigorously isolated. Evaluated on a unified hardware platform, AlgoPerf employs a diverse multi-task workload suite and a systematic optimizer comparison methodology to enable latency-accuracy co-evaluation across models, datasets, and hardware. Experiments reveal substantial latency disparities among mainstream optimizers, establish reproducible state-of-the-art baselines, and deliver the first quantitative, fair, and engineering-practical evaluation standard for training algorithm improvement.

Compare hyperparameter-tuned algorithms fairlyIdentify state-of-the-art training algorithms reliablyMeasure training time accurately and decide completion

Latest Papers

What's happening recently
View more

Towards Guided Descent: Optimization Algorithms for Training Neural Networks At Scale

Dec 20, 2025
AN
Ansh Nagwekar
🏛️ University of Pennsylvania

Stochastic Gradient Descent (SGD) and its variants lack rigorous theoretical foundations in over-parameterized neural networks, suffering from inefficient training and poor interpretability. Method: This paper proposes a principle-driven guided descent framework that unifies, for the first time, curvature-aware second-order approximations, layer-adaptive preconditioning (calibrated via condition number), and a dynamically parameterized maximum-update learning rate mechanism. It systematically elucidates the synergistic interplay between this framework and exponential moving average (EMA) as well as learning rate scheduling. Contribution/Results: The method achieves both scalability and theoretical interpretability while preserving training stability and significantly accelerating convergence—reducing large-model training time by an order of magnitude. Moreover, it enhances discriminative feature learning, simultaneously improving generalization performance and output consistency.

Bridges theoretical understanding with practical deployment strategiesExplores limitations of conventional methods like SGD on real-world dataInvestigates optimization algorithms for training neural networks at scale

This work addresses the diminished understanding of neural network fundamentals caused by the widespread use of high-level deep learning libraries. To bridge this gap, the authors construct a complete neural network framework from scratch, eschewing automatic differentiation and prebuilt modules. The implementation explicitly details forward and backward propagation, incorporates multiple activation functions, L2 regularization, and advanced optimizers such as Adam. Designed to balance pedagogical clarity with engineering scalability, the framework demonstrates numerical stability, correctness, and generalization capability on multiclass classification tasks. It thus provides a reproducible and extensible tool for both research and instruction, fostering deeper insight into the core principles of deep learning.

deep learning librarieseducational gapfundamental understanding

Traditional scaling laws rely solely on model and data scale to predict performance, neglecting other hyperparameters and thus struggling to achieve accurate prediction and efficient tuning under hardware constraints. This work proposes Configuration-to-Performance Scaling Laws (CPL), which, for the first time, incorporate the full training configuration into the modeling framework. By parameterizing this mapping with a large language model, the authors introduce a neuralized CPL (NCPL). Trained on open-source pretraining logs, NCPL enables joint optimization across multiple hyperparameters and predicts loss curves with 20–40% lower error than Chinchilla scaling laws. It generalizes effectively to regimes up to ten times the maximum compute budget observed in the training set and matches baseline methods in multi-hyperparameter tuning tasks.

hyperparameter tuninglarge language modelsperformance prediction

Hot Scholars

JK

Jan Kautz

Vice President of Research, NVIDIA Research
Computer VisionMachine LearningVisual Computing
SH

Siteng Huang

Alibaba DAMO Academy | ZJU | Westlake University
Vision-language ModelsGenerative ModelsEmbodied AI
XL

Xuyang Liu

Sichuan University
Vision-language ModelsModel CompressionToken CompressionTransfer Learning
ZY

Zihao Ye

NVIDIA, University of Washington
CompilersMachine Learning Systems
VG

Vinod Grover

Sr Distinguished Engineer, NVIDIA Corporation
Programming LanguagesCompilersDeep Learning