Score
Designs, implements, and analyzes numerical algorithms for constructing and manipulating M-matrices and block-structured matrices, producing componentwise-accurate factorizations, stable inversion methods, Schur complements, and singular-value computations. Builds block matrix algorithms and methods to compute matrix functions (for example the matrix square root) and matrix inverses with provable componentwise error bounds and numerical stability.
This work addresses the susceptibility of ill-conditioned M-matrices to subtractive cancellation in componentwise high-precision computations by introducing a novel approach based on a triplet representation. By integrating componentwise high-precision arithmetic with the structural properties of M-matrices, the authors present—for the first time—high-precision variants of the GTH algorithm, including non-blocking, recursive, and blocked formulations, all encapsulated within an object-oriented MATLAB interface. The resulting toolbox supports a range of operations such as linear system solving, LU factorization, Schur complement computation, matrix square roots, singular value decomposition, and solutions to nonsymmetric algebraic Riccati equations. Even under severe ill-conditioning, the framework preserves componentwise numerical accuracy, thereby substantially enhancing computational reliability.
This paper addresses the weak theoretical foundations of matrix decomposition in machine learning by systematically constructing a self-consistent, comprehensive, and modern-application-oriented pedagogical framework. Methodologically, it grounds the exposition in numerical linear algebra and matrix analysis, unifying classical decompositions—including LU, QR, SVD, and block triangular factorizations—while integrating numerical stability analysis and Hermitian/Hilbert space theory. Crucially, it bridges traditional numerical analysis with deep learning’s backpropagation setting, emphasizing differentiability and computational robustness of decompositions in algorithm design and model optimization. The primary contribution is a compact, dual-purpose (teaching and research) knowledge system that fills critical gaps in both theoretical coherence and machine-learning relevance present in existing literature, thereby providing rigorous mathematical foundations for high-dimensional data modeling and efficient training.
This work proposes a systematic framework to enhance the computational efficiency and numerical stability of evaluating high-degree matrix polynomials. Specifically, for polynomial degrees eight and higher, the method generates and validates stable coefficient sets that reduce the number of required matrix multiplications by one compared to the classical Paterson–Stockmeyer scheme. To address instability issues in the original formulation, the authors introduce structural variants and design a reliability metric to assess the expected numerical accuracy of candidate coefficient sets. Nonlinear polynomial systems are solved using variable-precision arithmetic (VPA), and an in-house tool, MatrixPolEval1, enables efficient screening and validation. Applied to matrix exponentials and geometric series, the approach achieves a saving of one matrix multiplication while maintaining comparable numerical accuracy.
To address the insufficient co-optimization of scalability and system stability in large-scale matrix computations on distributed-memory supercomputing platforms, this paper proposes the Block-Recursive Matrix Algorithm (BRMA) framework. BRMA uniquely integrates recursive block decomposition with distributed dynamic runtime control, incorporating task-graph–driven scheduling, a decentralized consensus protocol, and dynamic load rebalancing. Evaluated on a thousand-node cluster, BRMA achieves near-linear strong scaling, reduces communication overhead by 40%, and enables millisecond-scale fault recovery—significantly enhancing both scalability and fault tolerance. Its core contribution lies in unifying algorithmic scalability with system-level robustness, establishing a high-reliability, self-adaptive computational paradigm for exascale scientific computing.
Computing exact leverage scores for Kronecker-product structured matrices in large-scale least-squares problems is computationally prohibitive, while existing approximation methods incur statistical bias and high overhead. Method: We propose the first efficient exact leverage score algorithm tailored to Kronecker-structured matrices. Leveraging the inherent tensor structure, our method designs a near-linear-time framework for exact leverage score computation and sampling—bypassing costly full-matrix SVD or biased sketching approximations. Contribution/Results: Theoretically and empirically, our algorithm achieves significantly lower sampling error than state-of-the-art approximate methods (e.g., FJLT- or CountSketch-accelerated approaches), while maintaining substantially lower time complexity than full SVD. This work establishes the first scalable, exact, and efficient leverage score sampling scheme for Kronecker-structured matrices, enabling improved structured random projections and large-scale regression.
This work proposes a unified framework for solving systems of multivariate polynomials and rectangular multiparameter eigenvalue problems, implemented in the open-source MATLAB toolbox MacaulayLab. The approach leverages numerical linear algebra and Macaulay matrix constructions without relying on any specific polynomial basis or monomial ordering. It is the first method capable of efficiently handling both problem classes within a single framework while accurately characterizing positive-dimensional solution components at infinity. Numerical experiments demonstrate that the proposed method matches or surpasses the performance of established software packages such as PHCpack, PNLA, and MultiParEig. To support reproducible research, the authors provide an extensive suite of test cases alongside the toolbox.
This work addresses the gap between algorithmic prototypes and efficient implementations in scientific research by proposing a lightweight approach to translate statistical and machine learning algorithms—such as kernel ridge regression and stochastic gradient descent matrix factorization—from mathematical formulations into readable, high-performance C++ code. Leveraging the Eigen template library for core linear algebra operations—including kernel matrix construction, regularized solvers, and vectorized updates—the implementation seamlessly integrates into the Python ecosystem via pybind11, enabling efficient interoperability with NumPy arrays. The project provides concise, reproducible code examples that encapsulate common computational patterns in research, significantly lowering the barrier for researchers to adopt C++ for high-performance development while balancing performance, readability, and usability.