Score
Designs, implements, and analyzes data structures, algorithms, and numerical routines that represent, assemble, and operate on sparse matrices, vectors, and models; this includes implementing sparse formats and operations (e.g., SpMV), sparse matrix assembly, sparse linear solvers and Krylov methods, and sparse model training/inference such as sparse coding, regression, and representation learning. It also covers building sparse retrieval, routing, and selection algorithms and deriving complexity, numerical stability, approximation and error bounds (e.g., top-k approximation, sparse precision estimation) for these computations.
To address the low training efficiency and high computational/memory overhead of block-sparse models, this paper proposes an end-to-end differentiable training framework that abandons the conventional “dense-then-prune” paradigm, instead initializing directly from a sparse structure and dynamically optimizing block sizes during training. The method introduces three key components: gradient updates under structured sparsity constraints, block-level adaptive mask learning, and sparse-dense hybrid forward/backward propagation—ensuring both hardware compatibility and training stability. Experiments across multiple benchmarks demonstrate 40–65% reductions in computation and memory usage while matching the accuracy of dense baselines; moreover, the framework enables automatic block-size search. To our knowledge, this is the first structured sparse training approach that jointly optimizes block dimensions with model parameters during end-to-end training.
This work addresses load imbalance and memory bandwidth bottlenecks in sparse matrix–vector multiplication (SpMV) on parallel shared-memory systems. To mitigate these challenges, we propose six novel hybrid algorithms that jointly integrate dynamic task scheduling, NUMA-aware memory allocation, multithreaded optimization, and nonzero-element reordering. Furthermore, we systematically quantify the overhead–benefit trade-off of format conversion among common storage schemes (e.g., CSR, ELL, HYB), establishing— for the first time—a break-even threshold of 472 SpMV iterations. Our open-source implementation achieves an average 19% speedup over state-of-the-art methods across diverse multi-CPU architectures, significantly enhancing the efficiency of large-scale sparse computations.
Software verification of sparse matrix-vector multiplication (SpMV)—a core numerical kernel in high-performance computing (HPC)—lacks standardized, realistic benchmarks, hindering rigorous evaluation of verification tools. Method: This work systematically constructs the first SpMV verification benchmark, implemented within the PETSc framework. It includes rigorously reproduced serial and basic MPI-parallel versions that capture canonical sparse computation patterns used in scientific iterative solvers. An extensible, configurable verification testbed is introduced, supporting assertion injection, fault injection, and result comparison to enhance tool assessment capabilities. Contribution/Results: The open-source, fully reproducible benchmark has been adopted as a standard test suite by multiple verification tools, filling a critical gap in HPC software correctness validation. It establishes a foundational infrastructure for trustworthy software verification in scientific computing, enabling systematic evaluation and advancement of verification methodologies for sparse linear algebra kernels.
This work investigates the average-case computational complexity of sparse linear regression (SLR), focusing on whether polynomial-time algorithms exist for ill-conditioned design matrices—e.g., those with low rank or high correlation. The authors establish the first rigorous, instance-level reduction from classical worst-case lattice problems—specifically Bounded Distance Decoding (BDD)—to SLR. Their framework directly links the condition number of the lattice problem to the restricted eigenvalue condition of the SLR design matrix. This reduction holds in both identifiable and unidentifiable regimes. Leveraging worst-case-to-average-case hardness amplification, they prove that if BDD is hard in the worst case, then SLR remains computationally intractable on average for all polynomial-time algorithms. The result bridges a fundamental gap at the intersection of high-dimensional sparse statistics and computational complexity theory, providing the first evidence of average-case hardness for SLR under realistic design matrix conditions.
This work addresses the challenges of discovering differential equations, approximating functions, and estimating high-dimensional integrals from noisy, non-uniformly sampled data by introducing Sparse Orthogonal Regression Technique (SORT). The method reformulates equation discovery as a spectral coefficient learning problem, directly inferring coefficients of an orthogonal basis expansion from observational data via L1 regularization—without requiring a predefined symbolic library, explicit numerical integration, or inner product evaluations. Its core innovation lies in treating basis function design as the central modeling choice, enabling consistent model order scaling and multi-task reusability. Experiments demonstrate that when the basis functions align with the underlying problem structure, SORT matches or outperforms existing approaches under sparse sampling, noisy derivatives, and representation mismatch, while low-order dominant coefficients remain stable as model complexity increases.
This work presents the first systematic evaluation of Rust’s suitability for high-performance sparse linear algebra, addressing the longstanding trade-off between performance and memory safety in traditional scientific computing that relies on C/C++ and Fortran. The authors natively implement core operations—including sparse matrix-vector multiplication, the Lanczos Krylov method, and matrix exponential computation—leveraging compile-time monomorphization, SIMD vectorization, and careful FFI boundary analysis. Comprehensive benchmarks against established libraries such as Intel oneMKL, Eigen, PETSc, and PSBLAS demonstrate that Rust achieves performance on par with Eigen and PSBLAS in CSC format, approaching state-of-the-art levels while preserving memory safety. However, it still lags behind PETSc in block CSR optimizations, highlighting both the promise and current limitations of Rust for building efficient, safe numerical software stacks.