Score
Designs, implements, and analyzes algorithms and software for manipulating vectors and matrices, including dense-matrix operations, factorizations (LU, QR, Cholesky), linear solvers, and eigenvalue/SVD computations. Evaluates and optimizes their numerical correctness, stability, computational complexity, memory use, and performance (e.g., blocking, BLAS/LAPACK-style kernels, and parallelization) for dense linear-algebra workloads.
Scientific software deployment faces a productivity–performance gap between high-level prototyping environments (e.g., MATLAB) and production-grade C++ code: scripting languages sacrifice performance, while conventional C++ linear algebra libraries (e.g., BLAS/LAPACK) suffer from verbose interfaces and error-prone manual memory management. This paper introduces a novel C++ linear algebra library built on template metaprogramming and expression templates. It pioneers a compile-time expression optimization framework that enables operator fusion, lazy evaluation, and elimination of temporary objects—delivering MATLAB-like syntactic simplicity without explicit memory management. The design achieves both high code readability and performance competitive with hand-optimized C++. Benchmark results demonstrate speedups of several-fold over conventional C++ implementations. The library has been successfully deployed in production-scale machine learning and signal processing systems.
This work addresses the gap between algorithmic prototypes and efficient implementations in scientific research by proposing a lightweight approach to translate statistical and machine learning algorithms—such as kernel ridge regression and stochastic gradient descent matrix factorization—from mathematical formulations into readable, high-performance C++ code. Leveraging the Eigen template library for core linear algebra operations—including kernel matrix construction, regularized solvers, and vectorized updates—the implementation seamlessly integrates into the Python ecosystem via pybind11, enabling efficient interoperability with NumPy arrays. The project provides concise, reproducible code examples that encapsulate common computational patterns in research, significantly lowering the barrier for researchers to adopt C++ for high-performance development while balancing performance, readability, and usability.
This project addresses efficient randomized algorithms for three fundamental matrix computations under resource constraints: (1) low-rank positive semidefinite (PSD) matrix approximation, aiming for exact low-rank reconstruction with minimal element accesses; (2) estimation of implicit matrix properties—trace, diagonal entries, and row norms—using only matrix-vector products; and (3) numerically robust floating-point solutions to overdetermined linear least-squares problems. Methodologically, we propose a randomized pivoted Cholesky decomposition to drastically reduce element accesses in PSD approximation; introduce a novel leave-one-out framework for property estimation, achieving optimal sample complexity; and design the first randomized least-squares solver that simultaneously attains asymptotic acceleration—O(nd log d) time for n × d matrices—and numerical backward stability, matching the accuracy of direct methods. These contributions advance both theoretical guarantees and practical efficiency in large-scale, memory- or query-constrained numerical linear algebra.
This work addresses the efficient and numerically stable factorization of skew-symmetric matrices on shared-memory architectures, introducing a novel LTLᵀ decomposition (where L is unit lower triangular and T is tridiagonal). Methodologically, it is the first to systematically integrate formal derivation with classical algorithmic design; employs a multi-level computational fusion strategy to minimize memory traffic; and supports pivoting for numerical stability. Implementation-wise, it leverages the BLIS framework with hierarchical parallelism, cache-aware blocking, and fusion of Level-2 BLAS operations, realized concisely in C++ to ensure tight alignment between theory and code. Experiments demonstrate substantial speedups in both factorization and Pfaffian computation, achieving near-peak memory bandwidth utilization and resolving long-standing algorithmic and parallel implementation bottlenecks for skew-symmetric matrix factorization in dense linear algebra.
This work presents the first systematic evaluation of Rust’s suitability for high-performance sparse linear algebra, addressing the longstanding trade-off between performance and memory safety in traditional scientific computing that relies on C/C++ and Fortran. The authors natively implement core operations—including sparse matrix-vector multiplication, the Lanczos Krylov method, and matrix exponential computation—leveraging compile-time monomorphization, SIMD vectorization, and careful FFI boundary analysis. Comprehensive benchmarks against established libraries such as Intel oneMKL, Eigen, PETSc, and PSBLAS demonstrate that Rust achieves performance on par with Eigen and PSBLAS in CSC format, approaching state-of-the-art levels while preserving memory safety. However, it still lags behind PETSc in block CSR optimizations, highlighting both the promise and current limitations of Rust for building efficient, safe numerical software stacks.
This work addresses the inefficiencies of conventional GPU linear algebra libraries, which often fail to fully exploit hardware capabilities due to redundant memory accesses and runtime overhead. To overcome these limitations, the authors propose Bandicoot, a toolkit that leverages C++ template metaprogramming to perform expression fusion at compile time, thereby automatically generating highly optimized GPU kernels that saturate memory bandwidth—without relying on just-in-time compilation or runtime scheduling. Bandicoot provides an API compatible with Armadillo, facilitating straightforward migration of existing CPU codebases. Experimental results demonstrate that Bandicoot significantly outperforms PyTorch, TensorFlow, and JAX across multiple benchmarks, achieving substantial speedups in several scenarios.
This work addresses the memory bandwidth bottleneck encountered when performing QR decomposition of tall-and-skinny dense real matrices on GPUs. The authors systematically evaluate algorithms including Cholesky-QR2, SVQB, and Householder-based TSQR, and introduce two key optimizations: a Q-less QR strategy that avoids explicitly storing the orthogonal factor Q to reduce memory overhead, and a hybrid approach combining shared memory with a tree-based reduction structure to accelerate local computations. Experimental results on double-precision GPUs demonstrate that the optimized TSQR implementation significantly outperforms existing methods in memory- and compute-bound regimes and remains competitive with vendor-provided libraries, underscoring the critical role of specialized optimizations for QR decomposition of tall-and-skinny matrices.
This work addresses fundamental open problems in solving linear systems and computing eigenvalues by synthesizing consensus reached among experts in theoretical computer science and numerical analysis during a workshop at the Simons Institute. It presents the first systematic survey of interdisciplinary challenges, covering key techniques such as iterative solvers, low-rank approximations, and randomized sketching, while also extending to emerging areas including tensors, quantum systems, and matrix functions. The study distills these challenges into five major categories of critical questions, offering a clear research roadmap that advances the theory of computational complexity in linear algebra and guides the design of efficient algorithms, thereby fostering collaborative innovation between the two disciplines.
This work presents the first systematic evaluation and optimization of the Cerebras CS-3 platform for high-sparsity linear algebra computations, with a focus on sparse-dense matrix multiplication (SpMM) and sampled dense-dense matrix multiplication (SDDMM). Tailored sparse kernels are designed to align with the CS-3’s dataflow architecture, optimizing memory access patterns and I/O efficiency to enhance scalability for large-scale sparse matrices. Experimental results demonstrate that at 90% sparsity, the CS-3 achieves speedups of 100× and 20× over CPU baselines for SpMM and SDDMM, respectively. However, performance degrades significantly beyond 99% sparsity, falling below CPU performance. This study fills a critical gap in understanding the CS-3’s capabilities for sparse computation, revealing its substantial potential in moderately sparse regimes and identifying key bottlenecks under extreme sparsity.
本文探讨了通过生成式AI识别和利用线性代数中结构矩阵(如带状矩阵)的特性,以提高算法效率和内存使用,并提出了基于时间复杂度分析的评估标准。