Score
Designs and implements linear matrix preconditioners that transform parameters into a whitened coordinate space so that optimization and natural‑gradient updates behave isotropically. Builds and analyzes whitening preconditioner matrices to reconcile isotropic operations (e.g., isotropic additive noise) with anisotropic parameter scaling and to ensure stable preconditioning during noisy updates.
This work addresses the instability of existing matrix-sensing optimizers—such as Muon—whose momentum matrices exhibit pronounced coordinate-wise scale heterogeneity, rendering Newton–Schulz iterations highly sensitive to input conditioning. To resolve this, we propose Zeta, an optimizer that introduces a novel dual whitening mechanism with strict sequential dependency: it first applies coordinate whitening to rectify internal scale imbalances and then performs spectral whitening to enforce statistical isotropy, thereby yielding well-conditioned inputs for subsequent orthogonalization. This approach substantially reduces orthogonalization error and consistently outperforms strong baselines across diverse settings—including language models ranging from 0.6B to 8B parameters, mixture-of-experts architectures, and vision tasks—achieving faster convergence and improved generalization.
Stochastic Gradient Descent (SGD) often stagnates in late-stage training due to anisotropic curvature and gradient noise. This paper proposes a geometric optimization framework based on a preconditioner matrix (M), unifying the characterization of local condition number, lower bound on gradient noise, and basin stability of attraction. Under the (M)-induced Riemannian metric, the product of the effective condition number and the preconditioned noise level governs both convergence rate and steady-state error lower bound. Theoretically, this work provides the first basin stability guarantee—expressed in the (M)-norm—for non-convex landscapes. The method leverages symmetric positive-definite matrices to induce a Riemannian geometry, integrating stochastic optimization with local smoothness modeling, and supports both diagonal adaptive and curvature-aware preconditioner designs. Experiments on quadratic diagnostic tasks and three scientific machine learning benchmarks validate the predicted trade-off between convergence rate and noise amplification, demonstrating substantial improvements in convergence speed and training stability.
This work addresses the optimal design of diagonal preconditioners to simultaneously minimize both the classical worst-case condition number κ and the average-case ω-condition number, thereby accelerating convergence in preconditioned conjugate gradient (PCG) methods and optimization algorithms. We propose an affine-transformation-based pseudoconvex reformulation that converts the original nonconvex problem into an efficiently solvable form with only n-dimensional variables. Crucially, we establish for the first time that ω-optimal preconditioners inherently yield significantly reduced κ values, further enhancing PCG convergence. Instead of computationally expensive semidefinite programming (SDP), we employ a subgradient method for scalable and efficient optimization. Experiments demonstrate that our approach outperforms existing SDP-based methods in both scalability and computational efficiency, while delivering superior and more robust convergence acceleration across diverse problem instances.
Existing preconditioners for large-scale sparse linear systems suffer from low efficiency and poor generalization. Method: This paper proposes a novel learnable preconditioner that integrates algebraic preconditioning with graph neural networks (GNNs). It initializes the GNN with a classical ILU-type preconditioner and introduces a differentiable, condition-number-based loss function to explicitly optimize spectral properties during training. Additionally, it incorporates sparse structural priors and parameterized PDE modeling to ensure physical consistency and computational tractability. Results: On benchmark discretized parametric PDE systems, the method reduces iterative solver iterations by 30–50% compared to ILU and state-of-the-art neural preconditioners, achieves significantly improved condition numbers, and incurs only modest inference overhead. This work overcomes key limitations of purely data-driven and purely sparse-GNN-based preconditioners, establishing a new paradigm for interpretable, efficient, and generalizable learning in numerical linear algebra.
This work addresses efficiency bottlenecks in solving large-scale linear systems and approximating matrix norms. We propose a multilevel randomized sketching preconditioned iterative method, integrating Nyström low-rank approximation, sparse random sketching, and multilevel preconditioning. It establishes the first multilevel sketched preconditioning framework grounded in the natural average condition number. Theoretical contributions include: (1) optimal complexity $ ilde{O}(n^2 + d_lambda^omega)$ for solving regularized linear systems; (2) accelerated complexity $ ilde{O}(n^{2.065} + k^omega)$ for systems with $k$ outlying singular values; and (3) Schatten-$p$ norm approximation—particularly the nuclear norm—at $ ilde{O}(n^{2.11})$, improving upon the prior best $ ilde{O}(n^{2.18})$. These advances significantly enhance computational efficiency for key subproblems in applications such as Gaussian process regression.
本文通过介绍数值线性代数在偏微分方程、机器学习和数据同化中的应用,展示了如何使用少量核心概念解决大规模稀疏系统问题。
This work addresses the high computational cost of the KL-Shampoo optimizer in large language model pretraining by uncovering, for the first time, that its Kronecker factors exhibit a “spiky-flat” spectral structure. Building on this insight, the authors propose Pro-KLShampoo, which projects dominant directions of the Kronecker factors onto a low-dimensional subspace to preserve full spectral information, while approximating the remaining directions using shared eigenvalues and incorporating gradient momentum orthogonalization. This approach maintains algebraic equivalence to the original optimizer while substantially reducing both computational and memory overhead. Experiments across four model scales of GPT-2 and LLaMA demonstrate that Pro-KLShampoo consistently outperforms the original KL-Shampoo in terms of validation loss, peak memory usage per GPU, and convergence speed.
This work investigates the distinct error allocation behaviors of different Bregman divergences—Frobenius, von Neumann, and LogDet—in the spectral domain when the covariance matrix deviates from an exact Kronecker structure. Through spectral analysis of the covariance matrix, the study reveals a structural property wherein the top eigensubspace aligns closely with the Hessian, while the tail subspace is dominated by noise. Leveraging this insight, the authors propose a subspace-aware Kronecker optimizer that applies eigenvalue-based preconditioning in the reliable top subspace and introduces an adaptive isotropic acceleration constant in the noise-dominated bottom subspace. This approach effectively separates signal from noise components, yielding more stable and efficient optimization performance under non-ideal Kronecker structures.
This study addresses the challenge of adapting regularizers to data geometry in data-driven inverse problems by proposing an anisotropic dilation transform. This method adjusts radial statistics via a direction-preserving mapping, enabling optimal data adaptation and joint learning under a fixed base regularizer. Theoretically, leveraging the transitivity of the transform in radial statistical space, we derive a closed-form solution for optimal adaptation, prove its superiority over isotropic scaling, and establish finite-sample generalization bounds. Technically, the approach integrates Gibbs-type regularization, variational methods, and sparse ReLU network parameterization. Experiments demonstrate that the proposed method significantly enhances inverse problem performance in controlled two-dimensional settings and MNIST denoising tasks.
研究通过引入一个两参数家族来优化协方差矩阵,该方法能包含并扩展常见的度量方式,通过调整参数以改善特定问题的优化效果。