Score
Design and analyze regularization methods and optimizer modifications that promote or enforce approximate isometry (e.g., preserving singular values of input-output Jacobians) in parametric models, implemented as penalties or update rules applied separately from adaptive gradient steps so the adaptive optimizer does not cancel the isometry forces (Adam-style decoupling, Adamo). Build algorithms and update schedules that stabilize training dynamics and maintain dynamical isometry to improve optimization robustness and continued learning under non‑stationarity.
This work addresses a fundamental conflict in standard AdamW, where weight decay and adaptive gradient scaling engage in a “radial tug-of-war,” disrupting the learning of parameter directions and introducing noise. To resolve this, the authors propose AdamO, which explicitly decouples the radial (norm) and tangential (direction) dynamics of parameter updates. AdamO operates in orthogonal subspaces: it applies SGD-style updates to the radial component with curvature-adaptive step sizes, while employing Adam-style adaptive preconditioning for the directional component. Furthermore, it incorporates an architecture-aware update rule tailored for scale-invariant layers. Experiments demonstrate that AdamO consistently outperforms AdamW across vision and language tasks, achieving superior generalization and training stability without requiring additional constraints or hyperparameter tuning.
This work addresses the loss of plasticity in deep neural networks during continual learning, a critical limitation caused by non-stationary data distributions that hinders learning on subsequent tasks. For the first time, it explicitly links dynamic isometry to plasticity in continual learning and proposes a synergistic framework comprising an approximately dynamically isometric network architecture, an isometry-promoting regularization technique, and a decoupled optimizer named AdamO to jointly preserve model plasticity. Additionally, it introduces a novel mechanism to reactivate dormant ReLU units and reinterprets the limitations of existing methods through the lens of dynamic isometry. Evaluated across multiple supervised and reinforcement continual learning benchmarks, the proposed approach effectively mitigates plasticity loss and achieves performance on par with or superior to current state-of-the-art methods.
This work addresses the instability of temporal difference (TD) learning in offline reinforcement learning, where error amplification often leads to Q-value collapse. Viewing offline TD updates through the lens of control theory, the study models the learning process as a feedback system and reveals— for the first time—that the dynamic properties of the Adam optimizer can directly induce or suppress such collapse. To mitigate error propagation, the authors propose AdamO, a decoupled orthogonal correction mechanism incorporating a task-aligned budget constraint. Theoretical guarantees on worst-case task safety are established via spectral radius stability analysis and continuous-time dissipative dynamics. Empirically, AdamO significantly enhances both stability and performance across diverse offline RL benchmarks while maintaining broad compatibility with existing algorithms.
Existing adaptive optimization algorithms—such as AdaGrad, RMSProp, and Adam—lack a unified continuous-time characterization, hindering rigorous theoretical analysis and principled design. Method: The authors formulate these algorithms as first-order integro-differential equations, establishing the first unified continuous-time dynamical model that captures their implicit gradient-based dynamics and adaptive step-size mechanisms. Contribution/Results: Through rigorous numerical simulations, the proposed continuous model is shown to faithfully reproduce the discrete algorithms’ optimization trajectories, convergence rates, and adaptive behavior. This framework enables novel convergence analysis techniques and provides a theoretically grounded foundation for designing improved adaptive optimizers. By bridging discrete iterative updates with continuous dynamics, the work significantly extends the scope and rigor of continuous-time optimization theory.
This work addresses the convergence of RMSProp and Adam for generalized smooth nonconvex optimization. Under the weakest known assumptions—coordinate-wise generalized smoothness and affine noise variance—we establish the first tight theoretical guarantees. We introduce a novel descent lemma that overcomes critical challenges: adaptive step-size dependence, unbounded gradient estimates, and mismatched Lipschitz constants. Rigorously, we prove that both algorithms converge to an ε-stationary point in O(ε⁻⁴) iterations—the optimal rate matching the fundamental lower bound for nonconvex stochastic optimization. Our analysis operates under strictly weaker assumptions and yields tighter bounds than all prior works on RMSProp and Adam, thereby advancing the foundational understanding and reliability certification of adaptive optimization methods.
This work addresses the limitations of conventional optimizers, which entangle the magnitude and direction of weight updates, leading to unstable training dynamics that necessitate indirect stabilization techniques such as weight decay and learning rate warmup. To overcome this, the authors propose a Magnitude-Direction (MD) decoupling mechanism that explicitly decomposes each weight matrix—without altering model architecture—into a unit-norm directional component and learnable row- and column-wise magnitude gains. These components are optimized independently using separate learning rates, enabling precise control over magnitude and direction dynamics. The MD framework is compatible with any base optimizer (e.g., Adam, Muon) and eliminates reliance on weight decay or warmup schedules. Experiments demonstrate consistent improvements over carefully tuned baselines across diverse model scales, support learning rate transfer across model widths, and remain effective in large-scale Mixture-of-Experts (MoE) architectures.
This work addresses the lack of theoretical guarantees for the Adam algorithm in time-varying non-stationary systems, where existing analyses rely on the restrictive i.i.d. assumption. To bridge this gap, the authors develop a general theoretical framework tailored to dynamic environments by coupling the recursive dynamics of first- and second-order moments and introducing a novel stochastic Lyapunov function. They further establish analytical techniques for products of non-stationary dependent random matrices. Within this framework, they derive the first explicit bounds on both parameter tracking error and output prediction error for Adam, quantitatively characterizing the influence of step size, momentum parameters, gradient noise, and parameter drift. The theoretical findings are validated through experiments on both synthetic and real-world datasets, offering practical guidance for hyperparameter tuning.
本文解决了RMSprop优化器在全超参数控制下的收敛率问题,通过提供非渐近误差估计,并对所有允许的超参数值进行统一控制。
研究通过分析AdamW优化器中迷你批次的延迟影响,采用输入-状态-输出系统模型及线性化方法揭示了梯度扰动对未来损失的影响机制。
This work investigates whether higher-order adaptive Runge–Kutta (RK) optimizers genuinely improve neural network training and generalization under computationally matched conditions. We construct an Adam variant based on the Bogacki–Shampine 3(2) RK pair (RK-Adam) and conduct preregistered comparative experiments under a strict gradient computation budget. Our findings reveal that RK-Adam’s “adaptive” step size is effectively fixed, and its gradient averaging mechanism introduces an implicit regularization effect. Across all ten random seeds, RK-Adam outperforms learning-rate-matched Adam and AdamW but underperforms RMSprop and NAdam. Although it achieves approximately 40× lower training loss in full-batch settings, this does not translate into improved test accuracy. The study clarifies the actual behavior of RK-based optimizers and identifies the source of their limited generalization benefits.