scalable kernel matrix computation

Designs and implements algorithms, data structures, and software to compute, store, and manipulate large-scale kernel matrices efficiently, covering high-rank and Hamiltonian-structured kernels and custom data-processing kernel primitives. Builds memory-efficient batching, streaming I/O, and specialized linear-algebra routines and preprocessing pipelines to accelerate training and analysis workflows for large basis sets and high-rank representations.

scalablekernelmatrixcomputation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.3
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the gap between algorithmic prototypes and efficient implementations in scientific research by proposing a lightweight approach to translate statistical and machine learning algorithms—such as kernel ridge regression and stochastic gradient descent matrix factorization—from mathematical formulations into readable, high-performance C++ code. Leveraging the Eigen template library for core linear algebra operations—including kernel matrix construction, regularized solvers, and vectorized updates—the implementation seamlessly integrates into the Python ecosystem via pybind11, enabling efficient interoperability with NumPy arrays. The project provides concise, reproducible code examples that encapsulate common computational patterns in research, significantly lowering the barrier for researchers to adopt C++ for high-performance development while balancing performance, readability, and usability.

C++Eigenmachine learning

This work presents the first systematic evaluation of Rust’s suitability for high-performance sparse linear algebra, addressing the longstanding trade-off between performance and memory safety in traditional scientific computing that relies on C/C++ and Fortran. The authors natively implement core operations—including sparse matrix-vector multiplication, the Lanczos Krylov method, and matrix exponential computation—leveraging compile-time monomorphization, SIMD vectorization, and careful FFI boundary analysis. Comprehensive benchmarks against established libraries such as Intel oneMKL, Eigen, PETSc, and PSBLAS demonstrate that Rust achieves performance on par with Eigen and PSBLAS in CSC format, approaching state-of-the-art levels while preserving memory safety. However, it still lags behind PETSc in block CSR optimizations, highlighting both the promise and current limitations of Rust for building efficient, safe numerical software stacks.

memory safetyperformance evaluationRust

Even Faster Kernel Matrix Linear Algebra via Density Estimation

Oct 02, 2025
RS
Rikhav Shah
🏛️ MIT | UW-Madison

This work addresses key linear algebraic tasks—matrix-vector multiplication, matrix-matrix multiplication, spectral norm estimation, and all-entries summation—on $n imes n$ kernel matrices. We propose a unified approximation framework based on kernel density estimation (KDE), achieving $(1+varepsilon)$-relative error guarantees. Our core insight is to reduce each matrix operation to a set of KDE queries, thereby circumventing the standard $O(n^2)$ time barrier and attaining complexities nearly matching the optimal KDE runtime—for instance, $widetilde{O}(n/varepsilon^2)$ for all-entries summation. Theoretically, we establish the first conditional quadratic-time lower bounds for multiple kernel matrix problems and prove that our KDE-based paradigm achieves tight trade-offs between accuracy and efficiency. Empirically, our method significantly outperforms existing acceleration techniques on high-dimensional datasets.

Accelerating kernel matrix linear algebra operations via density estimationImproving runtime efficiency for matrix products and spectral normsReducing polynomial dependence on error and data size parameters

Giga-scale Kernel Matrix Vector Multiplication on GPU

Feb 02, 2022
RH
Robert Hu
🏛️ Amazon | University of Oxford | University of Adelaide | Université Paris Descartes

For kernel matrix-vector multiplication (KMVM) on billion-scale (10⁹), low-dimensional (D ≤ 7) datasets, conventional methods suffer from prohibitive O(N²) memory and time complexity. This work proposes F3M (Faster, Fast, and Free Memory Method)—the first GPU-accelerated KMVM algorithm empirically achieving linear time and memory complexity while maintaining approximation error below 10⁻³. F3M integrates three core techniques: GPU-optimized approximate kernel evaluation, synergistic low-rank decomposition and random projection for compression, and memory-aware tiling scheduling. On high-end GPUs, F3M computes billion-point KMVM in under 60 seconds. When integrated into FALKON, it delivers 1.5–5.5× speedup with <1% accuracy degradation. The method significantly enhances the scalability and practicality of kernel-based learning—particularly Gaussian process regression—for massive datasets.

Address scaling issues in KMVMImprove speed and efficiency on GPUPropose Faster-Fast and Free Memory Method

Distributed-memory Algorithms for Sparse Matrix Permutation, Extraction, and Assignment

Sep 25, 2025
EH
Elaheh Hassani
🏛️ Texas A&M University | Wake Forest University

Sparse matrix permutation, extraction, and assignment in distributed-memory systems suffer from high communication overhead and poor scalability. To address this, we propose an efficient parallel algorithm based on the “Identify-Exchange-Build” (IEB) paradigm. Our approach explicitly decouples communication and computation phases and employs synchronization-free multithreading to accelerate local submatrix construction, ensuring load balance while supporting complex use cases such as graph reordering and streaming graph processing. Experiments on heterogeneous supercomputing platforms—including Perlmutter—demonstrate that our method significantly outperforms CombBLAS and PETSc across diverse sparse matrix operations: it reduces communication volume by up to 37% and achieves strong scaling to over 10,000 cores. This work provides a highly scalable, low-overhead foundation for large-scale graph analytics and sparse linear algebra computations.

Accelerating local computations through synchronization-free multithreadingDeveloping scalable distributed algorithms for sparse matrix operationsReducing communication overhead compared to SpGEMM-based methods

Latest Papers

What's happening recently
View more

This work addresses the challenge of efficiently computing locality-driven integrals in quantum chemistry, whose matrix representations exhibit dynamic structural sparsity that is poorly handled by existing dense or generic sparse GPU algorithms. To overcome this, the authors propose KerneLDI, a GPU-oriented block-structured matrix multiplication framework that translates spatial locality into highly parallel block computations through a unified block filtering representation, custom dense block multiplication kernels, and co-optimized data layouts. Integrated with a multi-GPU parallelization strategy, KerneLDI achieves up to 10× acceleration in exchange-correlation energy evaluation while preserving numerical accuracy, significantly speeds up end-to-end self-consistent field iterations, and enhances ab initio molecular dynamics throughput by nearly sixfold.

GPU accelerationlocality-driven integrationmatrix multiplication

This study addresses the problem of reducing the additive complexity of $3\times3$ matrix multiplication. It proposes a rank-23 matrix multiplication kernel that combines linear programming reductions with sparse basis transformation search to optimize the computational structure, providing machine-verifiable certificates of correctness through exact coefficient expansion. By exploiting alternative bases, the method reduces the number of additions to 51, while achieving a low-complexity computation requiring only 56 additions in standard coordinates. This work presents the first rigorously and formally verified low-additive-complexity algorithm for $3\times3$ matrix multiplication, establishing a reliable benchmark and a novel paradigm for exploring theoretical lower bounds in this domain.

addition complexityalternative basesmatrix multiplication

This work addresses the scalability limitations of traditional kernel k-means clustering, which is constrained by single-GPU memory and struggles with million-scale datasets. To overcome this bottleneck, we propose a distributed kernel k-means algorithm designed for multi-GPU systems, reformulating core computations as communication-efficient distributed linear algebra primitives. Our approach introduces an innovative combination of 1.5D communication-avoiding partitioning and structured linear algebra techniques to minimize inter-GPU data movement. Evaluated on up to 256 GPUs, the method achieves 79.7% weak-scaling efficiency and a 4.2× speedup in strong scaling. Notably, clustering time is reduced from over one hour using a single-GPU sliding-window baseline to under two seconds, dramatically enhancing both scalability and performance for large-scale data clustering.

communication efficiencydistributed-memoryGPU memory limitation

This work addresses the lack of efficient solutions for computing a large batch of small-scale singular value decompositions (SVDs) on GPUs. The authors propose a GPU-accelerated batched SVD solver based on the one-sided Jacobi algorithm, co-designed with hardware architecture to exploit fine-grained parallelism, optimize memory access patterns, and support multiple floating-point precisions. Implemented on both NVIDIA and AMD GPU platforms, the solver demonstrates exceptional robustness and scalability across diverse matrix shapes, conditioning numbers, and precision configurations. Experimental results show that the proposed method significantly outperforms existing vendor-provided libraries and open-source solvers in terms of computational performance while maintaining numerical reliability.

batch SVDGPU computinghigh-performance computing

This work addresses the scalability limitations of Gaussian process regression (GPR) in high-dimensional settings, where computational complexity hinders application to large-scale data and high-dimensional inputs, particularly on incomplete grids. The authors propose CUTS-GPR, a novel method that integrates additive kernels, an incomplete tensor grid structure, and an efficient kernel matrix–vector multiplication algorithm to achieve, for the first time, exact GPR inference with near-linear or even linear time complexity. The approach enables rapid hyperparameter optimization and full posterior inference, completing the entire pipeline within hours on a dataset with N = 447,265 observations and D = 24 input dimensions. Demonstrated on high-dimensional potential energy surface modeling, CUTS-GPR substantially overcomes the longstanding scalability barrier of GPR, extending its applicability to problems involving billions of samples and thousands of features.

Gaussian process regressionhigh-dimensional dataincomplete grids

Hot Scholars

GC

Giuseppe Caire

Professor, Technical University of Berlin, Germany, and Professor of Electrical Engineering (on
Information TheoryCommunicationsSignal ProcessingStatistics
FX

François-Xavier Briol

Professor of Statistics and Machine Learning, UCL
Bayesian ComputationStatistical Machine LearningComputational StatisticsRobustness
HH

Haofeng Huang

Tsinghua University
Generative ModelsEfficient Machine LearningMachine Learning System