Score
Design and implement training pipelines that jointly optimize all components of a model from input to output by specifying differentiable architectures, losses, and optimization procedures so gradients propagate through the entire system. This includes constructing and balancing multiple loss terms, defining refinement and alignment objectives for intermediate representations, implementing backpropagation through composed modules or tools, and configuring optimizers and training schedules to converge the end-to-end model.
To address the challenge of efficiently optimizing high-dimensional configuration spaces in large-scale machine learning training, this paper proposes a scalable meta-gradient computation algorithm and the Smooth Model Training (SMT) framework—enabling, for the first time, end-to-end, differentiable joint optimization of training strategies. Methodologically, it integrates reverse-mode automatic differentiation through training loops, smooth modeling of training trajectories, and meta-gradient descent (MGD) to jointly optimize data selection, poisoning-resilient strategies, and learning rate scheduling. Key contributions are: (1) a breakthrough in scalable meta-gradient computation for large-scale training; and (2) the SMT framework, which ensures stability and convergence of MGD under realistic dynamic training conditions. Experiments demonstrate that the proposed data selection method significantly outperforms existing approaches; robustness against accuracy-degrading data poisoning attacks improves by an order of magnitude; and the fully automated learning rate scheduler matches or exceeds hand-crafted designs in performance.
Stochastic Gradient Descent (SGD) and its variants lack rigorous theoretical foundations in over-parameterized neural networks, suffering from inefficient training and poor interpretability. Method: This paper proposes a principle-driven guided descent framework that unifies, for the first time, curvature-aware second-order approximations, layer-adaptive preconditioning (calibrated via condition number), and a dynamically parameterized maximum-update learning rate mechanism. It systematically elucidates the synergistic interplay between this framework and exponential moving average (EMA) as well as learning rate scheduling. Contribution/Results: The method achieves both scalability and theoretical interpretability while preserving training stability and significantly accelerating convergence—reducing large-model training time by an order of magnitude. Moreover, it enhances discriminative feature learning, simultaneously improving generalization performance and output consistency.
This work addresses the limitation of conventional optimizers (e.g., Adam) that rely on hand-crafted gradient estimation heuristics. We propose Trainable Optimizer (TO), a framework that jointly trains a parameterized gradient estimator alongside the model parameters. Crucially, TO incorporates a pseudo-linear approximation of the estimator, enabling SGD-like convergence rates while substantially reducing gradient estimation variance. To enhance computational efficiency, we further introduce two lightweight variants requiring only minimal additional tensor operations. Theoretical analysis establishes convergence guarantees for both strongly convex and non-convex objectives. Empirical evaluation demonstrates that TO achieves faster convergence than Adam and other baselines across diverse benchmark tasks. Moreover, TO exhibits strong efficacy and scalability in fine-tuning large language models, validating its practical utility in modern deep learning settings.
Learned optimizers (L2Os) suffer from poor out-of-distribution generalization, limiting their applicability beyond the training data distribution. Method: This paper proposes a novel paradigm integrating classical optimization priors with data-driven modeling. It systematically incorporates fundamental optimization principles—specifically scale invariance and affine covariance—into the architecture design. We introduce a parameterized quasi-Newton update module explicitly constrained to preserve BFGS structure, and jointly optimize it via end-to-end training that unifies optimization-theoretic modeling, neural network architecture design, and meta-learning. Contribution/Results: The resulting enhanced BFGS algorithm significantly outperforms both standard L2Os and conventional solvers on unseen problem classes, dimensions, and condition numbers. It achieves over 40% improvement in cross-distribution generalization performance, establishing a new pathway toward more transferable and robust learned optimizers.
Existing learned optimizers (LOs) exhibit limited meta-generalization—particularly to unseen tasks requiring wider, deeper, or longer training trajectories. This work introduces μ-parameterization (μP) theory systematically into two mainstream LO architectures for the first time, proposing a μP-adapted lightweight meta-training paradigm. Methodologically, we derive theoretical scale-invariance conditions for LOs and design a low-overhead meta-training procedure (<250 GPU-hours). Experiments demonstrate that μLO matches or surpasses VeLO’s performance on large-width models—despite VeLO consuming 4,000 TPU-months—while improving meta-generalization in depth by 5× and extending maximal training-step generalization by 25×. This work establishes a rigorous theoretical foundation and an efficient implementation pathway for scalable, highly generalizable learned optimizers.
Existing hardware-software co-design tools struggle to accurately model memory consumption and backward-pass complexity in neural network training. This work proposes the first extension of the experimentally validated inference modeling framework, Stream, to the training domain, introducing a comprehensive framework for modeling and optimizing training on heterogeneous dataflow accelerators. The framework supports training workflow modeling, exploration of layer fusion configurations, and optimization of activation checkpointing strategies. Integrated with a genetic algorithm for hardware architecture search, it is validated on ResNet-18 and a small-scale GPT-2 model, effectively uncovering critical trade-offs between performance and memory in training-specific hardware design and identifying superior architectures and training strategies.
This work addresses the diminished understanding of neural network fundamentals caused by the widespread use of high-level deep learning libraries. To bridge this gap, the authors construct a complete neural network framework from scratch, eschewing automatic differentiation and prebuilt modules. The implementation explicitly details forward and backward propagation, incorporates multiple activation functions, L2 regularization, and advanced optimizers such as Adam. Designed to balance pedagogical clarity with engineering scalability, the framework demonstrates numerical stability, correctness, and generalization capability on multiclass classification tasks. It thus provides a reproducible and extensible tool for both research and instruction, fostering deeper insight into the core principles of deep learning.
Existing learned optimizers suffer from poor generalization and prohibitively high meta-training costs, hindering practical deployment. This work proposes a streamlined normalized optimizer architecture coupled with an enhanced meta-training strategy that drastically reduces computational overhead—requiring only 4.5 GPU hours—while remaining compatible with modern optimization techniques such as orthogonalization, layer-wise updates, and decoupled weight decay. The resulting learned optimizer scales robustly to billion-parameter models, outperforming prior methods on GPT-3 XL (1.3B) and demonstrating strong out-of-distribution generalization across diverse tasks, thereby overcoming limitations imposed by model scale and distributional shifts.
This work addresses the lack of non-convex convergence guarantees in PipeDream-style pipeline parallelism by proposing the Randomized PipeDream (RPD) framework, for which it establishes the first rigorous non-convex convergence theory. By introducing a randomized block SGD abstraction coupled with explicit modeling of communication delays, the analysis reveals that under steady-state conditions, the delay grows quadratically with the number of pipeline stages \(S\), leading to stale gradient terms scaling as \(\Theta(S^4)\). Empirical evaluations demonstrate that RPD outperforms LocalSGD in quadratic optimization and small-scale language model training, whereas LocalSGD exhibits superior performance as \(S\) increases in logistic regression tasks, highlighting a nuanced trade-off between the two methods across different problem settings.
This work proposes a differentiable programming–based framework for learning adaptive optimization algorithms to address the slow convergence and high per-iteration cost of traditional first-order methods in large-scale optimization. By embedding Fenchel–Rockafellar duality theory into automatic differentiation systems, the framework enables end-to-end training and adaptive refinement of duality-driven iterative schemes such as ADMM and PDHG. Implemented uniformly across major deep learning frameworks—including PyTorch, TensorFlow, and JAX—the approach significantly improves both computational efficiency and solution quality on a range of tasks, including linear programming, optimal power flow (OPF), Laplacian regularization, and neural network verification.