gradient-norm weighting

Designs and implements methods that compute gradient norms for individual loss terms and use those norms to adaptively scale or weight each objective during model training, thereby balancing learning speeds across tasks or views and stabilizing optimization. This includes building algorithms to compute per-loss or per-parameter gradient norms and integrating the resulting weights into the training/optimization loop.

gradient-normweighting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.55
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of conventional learning rate warmup strategies, which rely on heuristic hyperparameter tuning and lack theoretical grounding—particularly exhibiting instability under norm-constrained optimizers such as Muon and Lion. Building upon a generalized smoothness assumption that links local curvature to the suboptimality gap, the paper derives, for the first time, a learning rate schedule that naturally integrates both warmup and decay directly from convergence analysis. The resulting method is fully adaptive, requiring no additional hyperparameters and automatically adjusting warmup duration. Evaluated on LLaMA large language model pretraining, it consistently matches or surpasses the performance of manually tuned baselines across all experimental settings, significantly enhancing training efficiency and robustness.

adaptive schedulinglarge language modelslearning rate

Optimal Scaling Needs Optimal Norm

Oct 04, 2025
OF
Oleg Filatov
🏛️ Forschungszentrum Jülich

This work investigates the unified scaling规律 of optimal learning rate and batch size (η*, B*) under joint model and dataset scaling. We identify a novel “norm shift” phenomenon and establish that the operator norm of the output-layer weight matrix serves as a key invariant governing optimal hyperparameter selection, leading to a norm-guided scaling principle. Leveraging the Scion optimizer and Disco distributed training framework, we conduct over 2,000 experiments—spanning models up to 1.3B parameters and datasets up to 138B tokens—to empirically validate the universality of this norm-based condition: the output layer exhibits highest sensitivity, while hidden layers require comparatively lower learning rates. Crucially, we derive the first sufficient condition for optimal (η*, B*) pairs explicitly parameterized by dataset scale. This principle significantly enhances training stability and efficiency for large language models.

Determines optimal learning rate-batch size pairsEstablishes scaling rules for large language modelsIdentifies operator norm as scaling invariant

This work uncovers the mechanism by which the interaction between normalization and weight decay during deep neural network training triggers abrupt loss spikes: the scale invariance induced by normalization causes weight decay to continuously shrink the weight norm, leading to a sharp increase in loss landscape sharpness and optimization instability. To address this, the study introduces a novel concept—“weight norm criticality”—which elucidates why excessive weight decay, while improving generalization, inherently destabilizes training, and provides testable theoretical predictions. Through rigorous theoretical analysis and extensive experiments across multiple architectures, the authors establish a causal relationship between weight norm evolution and loss spikes, offering a principled foundation for balancing regularization strength and training stability.

loss spikesnormalizationtraining instability

This study investigates how effective learning rate drift induced by normalized updates affects training acceleration, stability, and resource scaling. Leveraging the random feature model and dynamical mean field theory (DMFT), the work elucidates the acceleration mechanisms and late-stage instabilities of normalized SGD, with theoretical predictions validated through linearized ResNet experiments. The primary contribution is the first solvable theoretical framework connecting normalization-induced acceleration, marginal stability, and width-batch allocation. Furthermore, this research quantifies the computational efficiency trade-offs between batch size and network width across distinct scaling regimes, empirically corroborating the predicted trends on CIFAR-5M.

adaptive learning rateedge of stabilitygradient normalization

MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters

Feb 04, 2024
AS
Arsalan Sharifnassab
🏛️ University of Alberta | Leiden University

To address the inefficiency and poor generalizability of manual hyperparameter tuning—particularly for learning rates—this paper proposes a dynamic online meta-optimization framework that formulates learning rate adaptation as a discounted cumulative regret minimization problem over time. The method employs a gradient-based meta-update mechanism, enabling plug-and-play integration with any first-order optimizer (e.g., SGD, Adam) to achieve decoupled, real-time, adaptive step-size optimization. Key contributions include: (i) the first formalization of meta-optimization as discounted regret minimization; and (ii) a low-complexity variant that preserves theoretical rigor while ensuring computational efficiency and strong generalization. Experiments across diverse tasks demonstrate faster convergence, enhanced robustness to initialization and task heterogeneity, competitive performance against hand-tuned optimal schedulers, and significantly lower computational overhead compared to conventional hyperparameter search methods.

Dynamically adjusting step sizes during model optimizationOptimizing meta-parameters for efficient machine learning trainingReducing regret by considering long-term impact of learning rates

Latest Papers

What's happening recently
View more

This study investigates whether gradient-free weight perturbation methods require perturbing all model parameters and what factors truly govern their effectiveness. Through controlled ablation experiments, the work systematically disentangles the effects of perturbation dimensionality, subspace choice, and perturbation norm, revealing for the first time that the perturbation norm—not the dimensionality or subspace—is the key determinant of performance. Empirical results demonstrate that once norms are matched, diverse subspaces—such as those derived from SVD bases or random frames—yield nearly identical outcomes. Remarkably, perturbing as few as 12–16 scalar parameters achieves 98.2% of the performance of full-parameter perturbation (an average gap of only 1.8 percentage points). This insight exhibits strong cross-model transferability and offers a novel perspective on efficient fine-tuning.

gradient-free adaptationlanguage modelsparameter-efficient

This work addresses the common degradation in perceptual image quality observed during reinforcement learning (RL) post-training for reward alignment, a phenomenon inadequately captured by existing reward proxies. The study identifies, for the first time, anomalous expansion of the velocity field norm as a key structural signature of this quality deterioration. To mitigate this issue, the authors propose NormGuard—a hinge-based penalty mechanism that activates only when the norm exceeds a predefined threshold, dynamically constraining its growth during training. Evaluated across two base models, three RL algorithms, and two reward functions, NormGuard consistently enhances image quality and forensic-level realism as assessed by multimodal large language models, with particularly pronounced gains in few-step inference settings. These improvements are not attributable to early stopping and are achieved without compromising reward performance.

flow-matchingnorm inflationperceptual quality

This work addresses the challenge in performative prediction where model deployment induces distributional shifts that complicate optimization. Existing approaches often rely on strong assumptions about the loss function and data distribution, limiting their applicability. To overcome this, the paper proposes a gradient-based adaptive optimization algorithm that explicitly estimates deployment-induced distribution shifts via finite differences, thereby accommodating a broader class of losses and distributions without stringent assumptions. The method supports high-dimensional optimization and incorporates a sample-efficient approximation strategy to reduce data requirements. Theoretical analysis establishes convergence guarantees for the proposed algorithm. Empirical results demonstrate that it converges faster and more stably than existing methods, exhibiting superior robustness and practicality across diverse experimental settings.

distribution shiftgradient-based methodsloss functions

This work proposes an adaptive scalar-step gradient descent method for non-convex optimization that overcomes the restrictive assumptions of traditional approaches. Conventional methods rely on strong regularity conditions—such as global Lipschitz or Hölder continuity of the full gradient—which lead to overly conservative step sizes. In contrast, the proposed algorithm leverages one-sided Hölder regularity to estimate local curvature along the descent direction and dynamically adjusts the step size via a sufficient decrease condition. This strategy relaxes the need for global regularity assumptions while permitting larger steps in flat regions without compromising convergence. Theoretically, the method guarantees optimal stationarity of iterates for non-convex objectives. Empirical results demonstrate superior performance over existing scalar-step gradient methods in binary classification and non-convex Hölder regression tasks, achieving lower final loss, smaller gradient norms, and wider classification margins.

adaptive gradient descentdescent directionnonconvex optimization

Hot Scholars

AG

Akshat Gupta

UC Berkeley
Knowledge EditingNatural Language ProcessingSpoken Language Modeling
ZX

Zi Xu

Professor of Mathematics, Shanghai University
Optimizationmathematical programming
KU

Kishor Upla

Associate Professor, Electronics Engineering Department, S.V. National Institute of Technology,
Signal & image processingCompressive sensing
SW

Shuche Wang

National University of Singapore
Information TheoryOptimizationMachine LearningSignal Processing
DL

Dongsheng Li

Professor, School of Computer Science, National University of Defense Technology
Distributed ComputingParallel ComputingCloud ComputingPeer-to-Peer Computing