direction-magnitude decomposition

Designs and implements parameterizations and algorithms that split vector-valued parameters or update vectors into a directional (unit-norm) component and a scalar magnitude (norm) component, and builds update rules or constraints that act separately on direction and on scale. Uses this decomposition to enable or analyze optimization with unknown or low effective rank, to improve numerical stability and convergence behavior, and to study dynamics such as transitions between saddle points.

direction-magnitudedecomposition

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.43
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the slow convergence often caused by improper rank selection in low-rank matrix optimization. To overcome this limitation, we propose a Direction-Magnitude Decomposition (DMD) framework that decouples direction and magnitude variables to enhance optimization efficiency, particularly in settings where the target rank is unknown. We introduce two novel variants—over-parameterized DMD and recursive DMD—that exploit the saddle-to-saddle dynamics observed along optimization trajectories, thereby significantly accelerating convergence. Theoretically, we establish that DMD achieves exponential acceleration over standard gradient descent. Empirical evaluations on matrix factorization, sensing, and completion tasks consistently demonstrate the superior efficiency of the proposed approach.

Burer-Monteiro formulationfactorization rank selectionlow-rank matrix optimization

This work addresses the challenge of characterizing how constraints shape the geometric structure of feasible regions in constrained optimization. The authors propose a geometric framework grounded in self-adjoint operators, which constructs local reachable subspaces by encoding computational or feasibility constraints and yields a pseudoinverse-weighted gradient as the optimal first-order update direction. This approach unifies the treatment of projection, spectral truncation, and multi-objective feasibility, revealing a constraint-induced warped ascent geometry. It establishes a compatibility principle between spectral compression and multi-objective structures, with the algorithm dynamically focusing on dominant spectral modes to provide a unified geometric characterization of gradient projection, spectral compression, and multi-objective feasible directions.

geometric mechanismmulti-objective compatibilityoptimization

This paper addresses the weak theoretical foundations of matrix decomposition in machine learning by systematically constructing a self-consistent, comprehensive, and modern-application-oriented pedagogical framework. Methodologically, it grounds the exposition in numerical linear algebra and matrix analysis, unifying classical decompositions—including LU, QR, SVD, and block triangular factorizations—while integrating numerical stability analysis and Hermitian/Hilbert space theory. Crucially, it bridges traditional numerical analysis with deep learning’s backpropagation setting, emphasizing differentiability and computational robustness of decompositions in algorithm design and model optimization. The primary contribution is a compact, dual-purpose (teaching and research) knowledge system that fills critical gaps in both theoretical coherence and machine-learning relevance present in existing literature, thereby providing rigorous mathematical foundations for high-dimensional data modeling and efficient training.

Cover limited scope of matrix decomposition analysisIntroduce matrix decomposition techniques and applicationsProvide mathematical tools for numerical linear algebra

This work addresses the challenge of integrating augmented Lagrangian and optimistic dual methods for equality-constrained optimization by proposing an additive hybrid framework that unifies matrix-valued augmentation and optimistic correction as distinct decompositions of a common correction matrix. By adaptively selecting the optimal splitting and stepsize through local spectral weighting, the method yields, for the first time, a closed-form hybrid update rule that jointly balances primal curvature and dual memory scale within a finite number of steps. Theoretical analysis reveals the equivalence and design flexibility between the two mechanisms, while experiments demonstrate that the proposed approach significantly outperforms individual strategies on nonlinear equality-constrained problems, achieving performance close to grid-search optimality and matching state-of-the-art first-order primal-dual algorithms under moderate ill-conditioning.

augmented Lagrangianconstrained optimizationfeasibility

In parametric dynamical systems, the Proper Orthogonal Decomposition (POD) basis drifts with parameters, degrading the accuracy of reduced-order models (ROMs). Method: This paper proposes the Projected Gaussian Process (pGP) framework—the first to formulate subspace adaptation as a statistical learning task mapping parameter space to the Grassmann manifold. It employs a two-stage geometric mapping: Euclidean space → horizontal space → Grassmann manifold, integrating POD, exponential/logarithmic maps, horizontal-space projection, and Gaussian process regression to enable uncertainty-aware POD subspace prediction while preserving manifold structure. Contribution/Results: Numerical experiments demonstrate that pGP significantly improves ROM accuracy and robustness in both parametric extrapolation and interpolation scenarios, and provides interpretable, calibrated confidence quantification—establishing a new paradigm for parameter-sensitive model reduction.

Adapting POD basis for parametric Reduced-Order ModelsMapping parameters to Grassmann manifold subspacesPredicting optimal subspaces using Gaussian Process regression

Latest Papers

What's happening recently
View more

This work addresses the limitations of conventional optimizers, which entangle the magnitude and direction of weight updates, leading to unstable training dynamics that necessitate indirect stabilization techniques such as weight decay and learning rate warmup. To overcome this, the authors propose a Magnitude-Direction (MD) decoupling mechanism that explicitly decomposes each weight matrix—without altering model architecture—into a unit-norm directional component and learnable row- and column-wise magnitude gains. These components are optimized independently using separate learning rates, enabling precise control over magnitude and direction dynamics. The MD framework is compatible with any base optimizer (e.g., Adam, Muon) and eliminates reliance on weight decay or warmup schedules. Experiments demonstrate consistent improvements over carefully tuned baselines across diverse model scales, support learning rate transfer across model widths, and remain effective in large-scale Mixture-of-Experts (MoE) architectures.

neural network trainingoptimizer couplingtraining stability

This work addresses the persistent challenge in process systems modeling of simultaneously achieving accuracy, simplicity, and physical interpretability—particularly in control applications where nonlinear expressiveness must be balanced against a preference for linear structures. The authors propose a convex hybrid modeling paradigm grounded in operator theory, which constrains models to interpretable subspaces or nonlinearly parameterized interpretable manifolds. By introducing a reparameterization technique based on “canonical features” in an augmented parameter space, the approach effectively integrates kernel methods with convex optimization. This framework enables the construction of kernel-based hybrid surrogate models over families of interpretable static and dynamic systems, significantly enhancing both predictive accuracy and computational efficiency while preserving physical interpretability across diverse process systems modeling scenarios.

convex learninghybrid modelinginterpretability

This work addresses optimization problems defined over products of simplices, such as low-rank learning of discrete multivariate probability distributions and function data registration based on the Square-Root Velocity Function (SRVF) representation. To tackle the inherent constraints, the authors propose a smooth reparameterization that is strictly convex element-wise, transforming the constrained problem into an unconstrained optimization over a Riemannian manifold. The resulting problem is solved via Riemannian gradient descent (RGD). Theoretical analysis shows that this reparameterization maps second-order KKT points on the manifold to weak second-order KKT points of the original problem, ensuring theoretical soundness while enhancing computational efficiency. Experiments demonstrate that RGD significantly outperforms projected gradient descent (PGD), achieving more accurate shape-preserving registration in functional data and efficiently solving probability tensor decomposition tasks.

functional data registrationoptimizationprobabilistic tensor decomposition

Hot Scholars

PX

Ping Xue

Department of Physics, Tsinghua University
Optics and Physics
TF

Tianfan Fu

Nanjing University
AI for DrugAI for ScienceLarge Language Model
HY

Hanchen Yang

Georgia Institute of Technology
Computer ArchitectureMachine Learning
HZ

Hao Zhao

Tsinghua University
Computer Vision
SZ

Shuigeng Zhou

Fudan University
DatabaseBioinformaticsMachine Learning