Score
Designs, builds, or analyzes iterative numerical algorithms that refine approximations using Newton–Raphson–style updates, including Newton reciprocal iteration and the shifted inverse method, to compute reciprocals or roots more accurately. Work covers proving and optimizing convergence and numerical stability, reducing precision and multiplication costs, enabling data-parallel multiplication steps, and adapting the iterations for integer/fixed‑point or secure‑arithmetic implementations.
This study addresses the challenge of simultaneously achieving numerical stability and computational efficiency in floating-point polynomial multiplication by proposing a novel algorithm with significantly simplified structure. The method integrates Newton iteration error analysis, divide-and-conquer strategies, and floating-point stability control theory to realize efficient and stable polynomial multiplication while preserving the optimal $O(np \log(np))$ time complexity and maintaining tight relative error bounds. Furthermore, the core ideas are extended to the near-convex Min-plus convolution problem, successfully optimizing the Bringmann–Cassis algorithm and yielding a logarithmic-factor speedup.
This work addresses the lack of efficient support for high-precision integer division in the range of $2^{15}$ to $2^{18}$ bits on general-purpose GPUs. We propose a fully integer-based Newton–Raphson division algorithm that leverages shift-based inversion, prefix sums, and multi-precision multiplication to enable data-parallel acceleration on CUDA platforms, achieving the first efficient implementation for this precision range. A performance cost model centered on the number of multiplications guides our algorithmic optimizations. Experimental results demonstrate that the proposed method attains near-theoretical-optimal performance within the target precision and significantly outperforms the state-of-the-art CGBN library in handling medium-sized large integer divisions.
High-order optimizers are rarely used in machine learning training due to their prohibitive computational overhead. Method: This paper systematically analyzes the impact of finite-precision arithmetic on Newton’s method convergence, establishing the first rigorous convergence theorems for mixed-precision variants—including quasi-Newton and inexact Newton methods—and enabling quantitative estimation of solution accuracy bounds. It further proposes the Generalized Gauss–Newton method GNₖ, which drastically reduces computational cost by computing only a subset of second-order derivatives while preserving full Newton-level performance on regression tasks. Results: Experiments demonstrate that GNₖ outperforms Adam on standard benchmarks while incurring significantly lower computational overhead than conventional second-order methods. This work advances practical deployment of high-order optimization by providing both theoretical guarantees and an efficient algorithmic design.
This work addresses the challenge of efficiently executing floating-point matrix multiplication (GEMM) on integer-tensor-accelerated hardware (e.g., NVIDIA A100/H100). We propose a precision-controllable integer tiling reconstruction method that decomposes floating-point GEMM into multiple exact integer matrix multiplications followed by floating-point accumulation. Our key contributions are: (i) the first lightweight error propagation model enabling precision-driven, adaptive estimation of the required number of integer tiles; and (ii) a systematic analysis revealing the dual impact of row/column scaling imbalance on both numerical accuracy and computational throughput. Experiments validate our theoretical error bounds and demonstrate accurate identification of precision failure boundaries under realistic imbalance scenarios. The method achieves floating-point-level accuracy while significantly improving integer hardware utilization, enabling explicit, tunable trade-offs between performance and precision.
This work addresses the computational inefficiency of matrix functions—such as square roots, inverse roots, and orthogonalization—in neural network training, which stems from traditional iterative methods’ reliance on prior spectral information and their inability to adapt to dynamically changing matrix spectra. The authors propose PRISM, a novel framework that enables adaptive computation of matrix functions without requiring any prior knowledge of the spectrum. PRISM constructs, at each iteration, a polynomial surrogate of the current spectrum using random sketching and relies predominantly on GPU-friendly matrix multiplications. This approach automatically adapts to spectral shifts during training, substantially reducing computational overhead. When integrated into Shampoo and Muon optimizers, PRISM maintains optimization accuracy while significantly decreasing both iteration counts and wall-clock runtime.
Traditional methods for computing elementary transcendental functions rely on differential equations or power series, often suffering from implementation complexity and high memory overhead. This work proposes the first unified fixed-point framework grounded in the Banach contraction mapping principle, modeling the exponential, trigonometric, and natural logarithm functions as the unique fixed points of doubling identities. By constructing residual-based strictly contractive iteration operators, the approach guarantees numerical stability and explicit convergence rates. The resulting floating-point kernel operates without lookup tables or memory accesses, achieving performance in throughput-constrained scenarios that matches or exceeds that of mainstream math libraries; notably, sine and cosine computations are consistently faster across all tested iteration depths.
This work proposes a systematic framework to enhance the computational efficiency and numerical stability of evaluating high-degree matrix polynomials. Specifically, for polynomial degrees eight and higher, the method generates and validates stable coefficient sets that reduce the number of required matrix multiplications by one compared to the classical Paterson–Stockmeyer scheme. To address instability issues in the original formulation, the authors introduce structural variants and design a reliability metric to assess the expected numerical accuracy of candidate coefficient sets. Nonlinear polynomial systems are solved using variable-precision arithmetic (VPA), and an in-house tool, MatrixPolEval1, enables efficient screening and validation. Applied to matrix exponentials and geometric series, the approach achieves a saving of one matrix multiplication while maintaining comparable numerical accuracy.
This work addresses the limitations of traditional iterative methods for solving large-scale systems of equations, including low efficiency, poor robustness, and difficulties in parameter selection. To overcome these challenges, it proposes a "hybrid iteration" paradigm that integrates classical numerical algorithms with machine learning. Building upon conventional iterative schemes such as Newton's method, this approach incorporates deep learning-based optimization strategies to enable adaptive hyperparameter tuning, thereby combining the flexibility of data-driven methods with the theoretical reliability and interpretability of traditional algorithms. Furthermore, this project systematically reviews state-of-the-art methodologies in this domain and identifies key open challenges. Ultimately, it establishes a clear research trajectory for developing efficient and robust solvers for scientific computing.
This work addresses the limitations of classical iterative methods, which rely on forward error and are constrained by the condition number of the matrix. It introduces a new paradigm using backward error as the convergence criterion. The key contributions include the first proof that Richardson iteration achieves a condition-number-independent $O(1/k)$ convergence rate in backward error for any positive semidefinite linear system. Building on this, the authors design an accelerated algorithm, MINBERR, attaining an $O(1/k^2)$ convergence rate. They further integrate backward error minimization into Krylov subspace methods and extend the approach to general linear systems. The resulting general-purpose solver has complexity $O(n^2/\varepsilon)$, while MINBERR achieves $O(n^2/\sqrt{\varepsilon})$, demonstrating superior numerical performance in benchmark experiments.
本文解决了计算矩阵的Lewis权重问题,通过固定点迭代法对所有p>2的情况提供了一种高精度计算方法。