train end-to-end models

Design and implement training pipelines that jointly optimize all components of a model from input to output by specifying differentiable architectures, losses, and optimization procedures so gradients propagate through the entire system. This includes constructing and balancing multiple loss terms, defining refinement and alignment objectives for intermediate representations, implementing backpropagation through composed modules or tools, and configuring optimizers and training schedules to converge the end-to-end model.

trainend-to-endmodels

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.6
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$241K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Optimizing ML Training with Metagradient Descent

Mar 17, 2025
LE
Logan Engstrom
🏛️ MIT | Stanford | UIUC

To address the challenge of efficiently optimizing high-dimensional configuration spaces in large-scale machine learning training, this paper proposes a scalable meta-gradient computation algorithm and the Smooth Model Training (SMT) framework—enabling, for the first time, end-to-end, differentiable joint optimization of training strategies. Methodologically, it integrates reverse-mode automatic differentiation through training loops, smooth modeling of training trajectories, and meta-gradient descent (MGD) to jointly optimize data selection, poisoning-resilient strategies, and learning rate scheduling. Key contributions are: (1) a breakthrough in scalable meta-gradient computation for large-scale training; and (2) the SMT framework, which ensures stability and convergence of MGD under realistic dynamic training conditions. Experiments demonstrate that the proposed data selection method significantly outperforms existing approaches; robustness against accuracy-degrading data poisoning attacks improves by an order of magnitude; and the fully automated learning rate scheduler matches or exceeds hand-crafted designs in performance.

Efficiently calculating metagradients for model trainingImproving dataset selection and learning rate schedulesOptimizing training setup for large-scale ML models

Towards Guided Descent: Optimization Algorithms for Training Neural Networks At Scale

Dec 20, 2025
AN
Ansh Nagwekar
🏛️ University of Pennsylvania

Stochastic Gradient Descent (SGD) and its variants lack rigorous theoretical foundations in over-parameterized neural networks, suffering from inefficient training and poor interpretability. Method: This paper proposes a principle-driven guided descent framework that unifies, for the first time, curvature-aware second-order approximations, layer-adaptive preconditioning (calibrated via condition number), and a dynamically parameterized maximum-update learning rate mechanism. It systematically elucidates the synergistic interplay between this framework and exponential moving average (EMA) as well as learning rate scheduling. Contribution/Results: The method achieves both scalability and theoretical interpretability while preserving training stability and significantly accelerating convergence—reducing large-model training time by an order of magnitude. Moreover, it enhances discriminative feature learning, simultaneously improving generalization performance and output consistency.

Bridges theoretical understanding with practical deployment strategiesExplores limitations of conventional methods like SGD on real-world dataInvestigates optimization algorithms for training neural networks at scale

A Trainable Optimizer

Aug 03, 2025
RW
Ruiqi Wang
🏛️ Northwestern University

This work addresses the limitation of conventional optimizers (e.g., Adam) that rely on hand-crafted gradient estimation heuristics. We propose Trainable Optimizer (TO), a framework that jointly trains a parameterized gradient estimator alongside the model parameters. Crucially, TO incorporates a pseudo-linear approximation of the estimator, enabling SGD-like convergence rates while substantially reducing gradient estimation variance. To enhance computational efficiency, we further introduce two lightweight variants requiring only minimal additional tensor operations. Theoretical analysis establishes convergence guarantees for both strongly convex and non-convex objectives. Empirical evaluation demonstrates that TO achieves faster convergence than Adam and other baselines across diverse benchmark tasks. Moreover, TO exhibits strong efficacy and scalability in fine-tuning large language models, validating its practical utility in modern deep learning settings.

Develop trainable optimizer replacing manual gradient methodsEnhance efficiency via simplified TO variants for faster convergenceProve pseudo-linear TO matches SGD convergence with less variance

From Learning to Optimize to Learning Optimization Algorithms

May 28, 2024
CC
Camille Castera
🏛️ University of Tübingen | Saarland University

Learned optimizers (L2Os) suffer from poor out-of-distribution generalization, limiting their applicability beyond the training data distribution. Method: This paper proposes a novel paradigm integrating classical optimization priors with data-driven modeling. It systematically incorporates fundamental optimization principles—specifically scale invariance and affine covariance—into the architecture design. We introduce a parameterized quasi-Newton update module explicitly constrained to preserve BFGS structure, and jointly optimize it via end-to-end training that unifies optimization-theoretic modeling, neural network architecture design, and meta-learning. Contribution/Results: The resulting enhanced BFGS algorithm significantly outperforms both standard L2Os and conventional solvers on unseen problem classes, dimensions, and condition numbers. It achieves over 40% improvement in cross-distribution generalization performance, establishing a new pathway toward more transferable and robust learned optimizers.

Designing learned optimization algorithms usable beyond training settingsDeveloping learning-enhanced BFGS algorithm adaptable to various test settingsSynergy between classical optimization and Learning to Optimize (L2O)

μLO: Compute-Efficient Meta-Generalization of Learned Optimizers

May 31, 2024
BT
Benjamin Thérien
🏛️ Université de Montréal | Mila – Quebec AI Institute | Samsung | Flatiron Institute | Concordia University

Existing learned optimizers (LOs) exhibit limited meta-generalization—particularly to unseen tasks requiring wider, deeper, or longer training trajectories. This work introduces μ-parameterization (μP) theory systematically into two mainstream LO architectures for the first time, proposing a μP-adapted lightweight meta-training paradigm. Methodologically, we derive theoretical scale-invariance conditions for LOs and design a low-overhead meta-training procedure (<250 GPU-hours). Experiments demonstrate that μLO matches or surpasses VeLO’s performance on large-width models—despite VeLO consuming 4,000 TPU-months—while improving meta-generalization in depth by 5× and extending maximal training-step generalization by 25×. This work establishes a rigorous theoretical foundation and an efficient implementation pathway for scalable, highly generalizable learned optimizers.

Address optimization challenges in wider networks than meta-trainedEnhance generalization to deeper networks and longer training horizonsImprove meta-generalization of learned optimizers for unseen tasks

Latest Papers

What's happening recently
View more

Existing hardware-software co-design tools struggle to accurately model memory consumption and backward-pass complexity in neural network training. This work proposes the first extension of the experimentally validated inference modeling framework, Stream, to the training domain, introducing a comprehensive framework for modeling and optimizing training on heterogeneous dataflow accelerators. The framework supports training workflow modeling, exploration of layer fusion configurations, and optimization of activation checkpointing strategies. Integrated with a genetic algorithm for hardware architecture search, it is validated on ResNet-18 and a small-scale GPT-2 model, effectively uncovering critical trade-offs between performance and memory in training-specific hardware design and identifying superior architectures and training strategies.

backpropagation complexityhardware-software co-designheterogeneous accelerators

This work addresses the diminished understanding of neural network fundamentals caused by the widespread use of high-level deep learning libraries. To bridge this gap, the authors construct a complete neural network framework from scratch, eschewing automatic differentiation and prebuilt modules. The implementation explicitly details forward and backward propagation, incorporates multiple activation functions, L2 regularization, and advanced optimizers such as Adam. Designed to balance pedagogical clarity with engineering scalability, the framework demonstrates numerical stability, correctness, and generalization capability on multiclass classification tasks. It thus provides a reproducible and extensible tool for both research and instruction, fostering deeper insight into the core principles of deep learning.

deep learning librarieseducational gapfundamental understanding

Existing learned optimizers suffer from poor generalization and prohibitively high meta-training costs, hindering practical deployment. This work proposes a streamlined normalized optimizer architecture coupled with an enhanced meta-training strategy that drastically reduces computational overhead—requiring only 4.5 GPU hours—while remaining compatible with modern optimization techniques such as orthogonalization, layer-wise updates, and decoupled weight decay. The resulting learned optimizer scales robustly to billion-parameter models, outperforming prior methods on GPT-3 XL (1.3B) and demonstrating strong out-of-distribution generalization across diverse tasks, thereby overcoming limitations imposed by model scale and distributional shifts.

learned optimizersmeta-generalizationmeta-training cost

This work addresses the lack of non-convex convergence guarantees in PipeDream-style pipeline parallelism by proposing the Randomized PipeDream (RPD) framework, for which it establishes the first rigorous non-convex convergence theory. By introducing a randomized block SGD abstraction coupled with explicit modeling of communication delays, the analysis reveals that under steady-state conditions, the delay grows quadratically with the number of pipeline stages \(S\), leading to stale gradient terms scaling as \(\Theta(S^4)\). Empirical evaluations demonstrate that RPD outperforms LocalSGD in quadratic optimization and small-scale language model training, whereas LocalSGD exhibits superior performance as \(S\) increases in logistic regression tasks, highlighting a nuanced trade-off between the two methods across different problem settings.

convergencedistributed trainingmodel parallelism

This work proposes a differentiable programming–based framework for learning adaptive optimization algorithms to address the slow convergence and high per-iteration cost of traditional first-order methods in large-scale optimization. By embedding Fenchel–Rockafellar duality theory into automatic differentiation systems, the framework enables end-to-end training and adaptive refinement of duality-driven iterative schemes such as ADMM and PDHG. Implemented uniformly across major deep learning frameworks—including PyTorch, TensorFlow, and JAX—the approach significantly improves both computational efficiency and solution quality on a range of tasks, including linear programming, optimal power flow (OPF), Laplacian regularization, and neural network verification.

differentiable programmingfirst-order methodslarge-scale problems

Hot Scholars

AY

Aylin Yener

Roy and Lois Chope Professor, The Ohio State University
green communicationsphysical layer securitysemantic communications6G
DS

Davide Scaramuzza

Professor of Robotics and Perception, University of Zurich
RoboticsRobot VisionMicro Air VehiclesSLAM
WZ

Wenjun Zhang

City University of Hong Kong
Thin film technologynanomaterials and nanodevices
QY

Qianqian Yang

Zhejiang University
Information TheoryWireless AISemantic CommunicationMachine Learning
ZS

Zhiguo Shi

IEEE Fellow, IET Fellow, Qiushi Distinguished Professor, Zhejiang University
Statistic signal processingDrone SurveillanceInternet of ThingsSystem Security