apply linear algebra

Designs, implements, and analyzes algorithms and software for manipulating vectors and matrices, including dense-matrix operations, factorizations (LU, QR, Cholesky), linear solvers, and eigenvalue/SVD computations. Evaluates and optimizes their numerical correctness, stability, computational complexity, memory use, and performance (e.g., blocking, BLAS/LAPACK-style kernels, and parallelization) for dense linear-algebra workloads.

applylinearalgebra

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.95
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$203K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Armadillo: An Efficient Framework for Numerical Linear Algebra

Feb 05, 2025
CS
Conrad Sanderson
🏛️ Griffith University | NumFOCUS Inc. | CSIRO

Scientific software deployment faces a productivity–performance gap between high-level prototyping environments (e.g., MATLAB) and production-grade C++ code: scripting languages sacrifice performance, while conventional C++ linear algebra libraries (e.g., BLAS/LAPACK) suffer from verbose interfaces and error-prone manual memory management. This paper introduces a novel C++ linear algebra library built on template metaprogramming and expression templates. It pioneers a compile-time expression optimization framework that enables operator fusion, lazy evaluation, and elimination of temporary objects—delivering MATLAB-like syntactic simplicity without explicit memory management. The design achieves both high code readability and performance competitive with hand-optimized C++. Benchmark results demonstrate speedups of several-fold over conventional C++ implementations. The library has been successfully deployed in production-scale machine learning and signal processing systems.

Adapting research prototypes to production-grade codeEfficient linear algebra for machine learning applicationsSimplifying use of BLAS and LAPACK libraries

This work addresses the gap between algorithmic prototypes and efficient implementations in scientific research by proposing a lightweight approach to translate statistical and machine learning algorithms—such as kernel ridge regression and stochastic gradient descent matrix factorization—from mathematical formulations into readable, high-performance C++ code. Leveraging the Eigen template library for core linear algebra operations—including kernel matrix construction, regularized solvers, and vectorized updates—the implementation seamlessly integrates into the Python ecosystem via pybind11, enabling efficient interoperability with NumPy arrays. The project provides concise, reproducible code examples that encapsulate common computational patterns in research, significantly lowering the barrier for researchers to adopt C++ for high-performance development while balancing performance, readability, and usability.

C++Eigenmachine learning

Make the most of what you have: Resource-efficient randomized algorithms for matrix computations

Dec 17, 2025
EN
Ethan N. Epperly
🏛️ California Institute of Technology

This project addresses efficient randomized algorithms for three fundamental matrix computations under resource constraints: (1) low-rank positive semidefinite (PSD) matrix approximation, aiming for exact low-rank reconstruction with minimal element accesses; (2) estimation of implicit matrix properties—trace, diagonal entries, and row norms—using only matrix-vector products; and (3) numerically robust floating-point solutions to overdetermined linear least-squares problems. Methodologically, we propose a randomized pivoted Cholesky decomposition to drastically reduce element accesses in PSD approximation; introduce a novel leave-one-out framework for property estimation, achieving optimal sample complexity; and design the first randomized least-squares solver that simultaneously attains asymptotic acceleration—O(nd log d) time for n × d matrices—and numerical backward stability, matching the accuracy of direct methods. These contributions advance both theoretical guarantees and practical efficiency in large-scale, memory- or query-constrained numerical linear algebra.

Designing resource-efficient randomized algorithms for matrix computations.Developing accurate low-rank approximations with minimal matrix entry access.Ensuring backward stability in randomized linear least squares algorithms.

Skew-Symmetric Matrix Decompositions on Shared-Memory Architectures

Nov 15, 2024
IS
Ishna Satyarth
🏛️ Southern Methodist University | NVIDIA

This work addresses the efficient and numerically stable factorization of skew-symmetric matrices on shared-memory architectures, introducing a novel LTLᵀ decomposition (where L is unit lower triangular and T is tridiagonal). Methodologically, it is the first to systematically integrate formal derivation with classical algorithmic design; employs a multi-level computational fusion strategy to minimize memory traffic; and supports pivoting for numerical stability. Implementation-wise, it leverages the BLIS framework with hierarchical parallelism, cache-aware blocking, and fusion of Level-2 BLAS operations, realized concisely in C++ to ensure tight alignment between theory and code. Experiments demonstrate substantial speedups in both factorization and Pfaffian computation, achieving near-peak memory bandwidth utilization and resolving long-standing algorithmic and parallel implementation bottlenecks for skew-symmetric matrix factorization in dense linear algebra.

Enhancing parallel CPU performance via fused operations and BLISFactorizing skew-symmetric matrices into LTL^T decompositionImproving numerical stability with pivoting for matrix operations

This work presents the first systematic evaluation of Rust’s suitability for high-performance sparse linear algebra, addressing the longstanding trade-off between performance and memory safety in traditional scientific computing that relies on C/C++ and Fortran. The authors natively implement core operations—including sparse matrix-vector multiplication, the Lanczos Krylov method, and matrix exponential computation—leveraging compile-time monomorphization, SIMD vectorization, and careful FFI boundary analysis. Comprehensive benchmarks against established libraries such as Intel oneMKL, Eigen, PETSc, and PSBLAS demonstrate that Rust achieves performance on par with Eigen and PSBLAS in CSC format, approaching state-of-the-art levels while preserving memory safety. However, it still lags behind PETSc in block CSR optimizations, highlighting both the promise and current limitations of Rust for building efficient, safe numerical software stacks.

memory safetyperformance evaluationRust

Latest Papers

What's happening recently
View more

This work addresses the inefficiencies of conventional GPU linear algebra libraries, which often fail to fully exploit hardware capabilities due to redundant memory accesses and runtime overhead. To overcome these limitations, the authors propose Bandicoot, a toolkit that leverages C++ template metaprogramming to perform expression fusion at compile time, thereby automatically generating highly optimized GPU kernels that saturate memory bandwidth—without relying on just-in-time compilation or runtime scheduling. Bandicoot provides an API compatible with Armadillo, facilitating straightforward migration of existing CPU codebases. Experimental results demonstrate that Bandicoot significantly outperforms PyTorch, TensorFlow, and JAX across multiple benchmarks, achieving substantial speedups in several scenarios.

compile-time fusionease of useGPU linear algebra

This work addresses the memory bandwidth bottleneck encountered when performing QR decomposition of tall-and-skinny dense real matrices on GPUs. The authors systematically evaluate algorithms including Cholesky-QR2, SVQB, and Householder-based TSQR, and introduce two key optimizations: a Q-less QR strategy that avoids explicitly storing the orthogonal factor Q to reduce memory overhead, and a hybrid approach combining shared memory with a tree-based reduction structure to accelerate local computations. Experimental results on double-precision GPUs demonstrate that the optimized TSQR implementation significantly outperforms existing methods in memory- and compute-bound regimes and remains competitive with vendor-provided libraries, underscoring the critical role of specialized optimizations for QR decomposition of tall-and-skinny matrices.

dense matrix decompositionGPU computingmemory-bandwidth limited

This work addresses fundamental open problems in solving linear systems and computing eigenvalues by synthesizing consensus reached among experts in theoretical computer science and numerical analysis during a workshop at the Simons Institute. It presents the first systematic survey of interdisciplinary challenges, covering key techniques such as iterative solvers, low-rank approximations, and randomized sketching, while also extending to emerging areas including tensors, quantum systems, and matrix functions. The study distills these challenges into five major categories of critical questions, offering a clear research roadmap that advances the theory of computational complexity in linear algebra and guides the design of efficient algorithms, thereby fostering collaborative innovation between the two disciplines.

eigenvalue problemslinear systemsmatrix computations

This work presents the first systematic evaluation and optimization of the Cerebras CS-3 platform for high-sparsity linear algebra computations, with a focus on sparse-dense matrix multiplication (SpMM) and sampled dense-dense matrix multiplication (SDDMM). Tailored sparse kernels are designed to align with the CS-3’s dataflow architecture, optimizing memory access patterns and I/O efficiency to enhance scalability for large-scale sparse matrices. Experimental results demonstrate that at 90% sparsity, the CS-3 achieves speedups of 100× and 20× over CPU baselines for SpMM and SDDMM, respectively. However, performance degrades significantly beyond 99% sparsity, falling below CPU performance. This study fills a critical gap in understanding the CS-3’s capabilities for sparse computation, revealing its substantial potential in moderately sparse regimes and identifying key bottlenecks under extreme sparsity.

AI acceleratorCerebras CS-3SDDMM

Hot Scholars

YL

Yuhang Liu

The University of Adelaide
Representation LearningLLMsLatent Variable ModelsResponsible AI
NH

Nhat Ho

Assistant Professor at University of Texas, Austin
Machine LearningBayesian StatisticsOptimizationOptimal Transport
YK

Yoon Kim

Associate Professor, MIT
Machine LearningNatural Language ProcessingDeep Learning
SL

Simon Lucey

Director, Australian Institute for Machine Learning (AIML) + CommBank AI Scholar
Computer VisionMachine LearningArtificial Intelligence
HS

Hemanth Saratchandran

Australian Institute for Machine Learning/Adelaide University + CommBank AI Scholar
MathematicsMachine Learning