Score
Designs and implements algorithms that use Cholesky factorization to compute the action of a matrix inverse (inverse-vector products) without forming the dense inverse, including implicit Cholesky and inverse-free evaluation methods. Builds and analyzes implementations that preserve sparse-memory bounds (e.g., O(nnz(L)) or better) and support exact or high-precision arithmetic requirements while avoiding explicit inverse formation.
This study addresses the bottleneck of sparse Cholesky solvers struggling to accommodate mixed sparsity patterns due to fixed data structures by proposing a structure-adaptive tiled Cholesky factorization method. A lightweight selector dynamically routes computations to dense, semi-sparse, or sparse paths, while a novel “active-column tile” data structure enables semi-sparse patterns to achieve BLAS-3 efficiency. Combined with cross-matrix and tile-level adaptive storage alongside GPU acceleration, the approach efficiently supports repeated factorization scenarios in INLA. Experimental results demonstrate that on CPUs, the proposed method outperforms the best fixed strategies by 1.6–2.6× and existing solvers by 1.8–12.5×. Furthermore, the GPU implementation achieves an additional 1.2–6.3× speedup over CPU nodes.
Cholesky decomposition of symmetric positive-definite matrices in Gaussian process inference suffers from numerical instability and computational inefficiency. Method: We propose a novel pivoted Cholesky pivoting strategy that integrates entropy maximization from Bayesian nonparametric inference with information gain principles from active learning, enabling an online-updatable, low-overhead diagonal adaptation mechanism. This strategy is synergistically combined with preconditioned iterative solvers and sparse regression frameworks. Contribution/Results: The proposed method significantly improves both uncertainty quantification accuracy and computational speed in sparse Gaussian process inference. Empirical evaluation across diverse benchmark tasks demonstrates consistent superiority over standard pivoted Cholesky baselines, with negligible additional computational overhead. The approach is theoretically grounded in probabilistic inference principles and exhibits strong practical applicability in large-scale GP modeling.
This paper addresses the challenge of balancing computational efficiency and approximation accuracy in sparse approximate inverse Cholesky decomposition of high-dimensional kernel matrices. We propose a dynamic sparsity pattern construction method based on greedy conditional mutual information maximization—marking the first integration of mutual information maximization into sparse structure learning, enabling data-adaptive pivot selection. Unlike conventional geometric neighborhood constraints, our approach jointly leverages partial Cholesky updates and KL-divergence minimization to support efficient multi-objective aggregation during decomposition. Theoretical analysis shows the time complexity reduces to *O(Nk²)*, substantially improving upon the *O(N³)* cost of standard Cholesky decomposition. Experiments on Gaussian process regression and image classification demonstrate superior preprocessing accuracy and faster conjugate gradient convergence compared to k-nearest-neighbor and other baseline methods. Our framework establishes a scalable, high-fidelity sparse approximation paradigm for large-scale kernel learning.
Recursive matrix multiplication faces a fundamental trade-off between numerical accuracy and computational efficiency. Method: We propose a unified framework for automatically constructing high-accuracy, low-complexity recursive bilinear algorithms. First, we unify and extend norm-based error bounds—covering Frobenius, spectral, and induced norms—to arbitrary recursive bilinear schemes. Second, we introduce a three-stage optimization pipeline: (i) heuristic search minimizing operand count in bilinear formulations; (ii) growth-factor minimization over tensor decomposition orbits; and (iii) sparsification via alternative bases. Contribution/Results: Our framework yields a novel non-commutative 2×2 matrix multiplication algorithm requiring only seven scalar multiplications—matching Strassen’s asymptotic complexity O(n^log₂7) while substantially improving numerical stability. The approach generalizes to established recursive algorithms including those of Smirnov et al., demonstrating broad applicability and enabling scalable, automated generation of high-accuracy fast matrix multiplication schemes.
This study addresses the issues of excessive fill-in in approximate Cholesky factorization and the lack of theoretical guarantees and graph connectivity in existing sampling rules. To overcome these limitations, this work proposes a volume-sampling-based elimination strategy that optimizes the elimination ordering by uniformly and randomly generating maximum spanning trees of product cliques. This approach effectively integrates the theoretical correctness of Kyng’s framework with the practical connectivity of Gao’s method, strictly preserving marginal edge distributions while ensuring graph connectivity. By combining linear time complexity with O(log n)-depth parallel computing techniques, the proposed algorithm yields a concise, provably correct, and efficiently parallelizable scheme for approximate Cholesky factorization.
本文探讨了通过生成式AI识别和利用线性代数中结构矩阵(如带状矩阵)的特性,以提高算法效率和内存使用,并提出了基于时间复杂度分析的评估标准。
该论文提出了一种使用48次乘法和216次其他运算的快速4x4矩阵乘法算法,并通过递归应用及基变换优化,降低了计算复杂度。
本文通过理论分析解决了随机旋转Cholesky算法在大正半定矩阵低秩逼近中的误差估计问题,证明了该方法能在接近最优的复杂度内达到良好的逼近效果。
This work proposes a fast recursive framework for efficiently computing the truncated Neumann series $S_k(A) = I + A + \cdots + A^{k-1}$ by leveraging high-radix kernel functions $T_m(B) = I + B + \cdots + B^{m-1}$, substantially reducing the number of required matrix multiplications. The key contributions include the first exact radix-9 rational-coefficient three-product kernel and a residual-based recursive structure that enables stable application of approximate kernels. Furthermore, a radix-15 kernel is designed, achieving the best-known asymptotic complexity to date. Theoretical analysis shows that the radix-9 and radix-15 kernels reduce the multiplication count to $1.58\log_2 k$ and $1.54\log_2 k$, respectively, yielding approximately 21% speedup over conventional repeated squaring. Experimental results corroborate the predicted acceleration.
This work addresses the underutilization of CPU resources in modern heterogeneous high-performance computing (HPC) systems when solving large-scale symmetric positive-definite linear systems using GPU-only approaches. Leveraging the SYCL programming model, the authors present the first heterogeneous implementations of the conjugate gradient (CG) method and Cholesky decomposition that operate across multi-vendor CPU-GPU platforms, including NVIDIA, AMD, and Intel architectures. Experimental results demonstrate that the heterogeneous CG solver achieves up to 32% speedup over GPU-only execution on large matrices, while the heterogeneous Cholesky decomposition attains a 29% acceleration. Furthermore, across diverse hardware vendors, the Cholesky solver consistently delivers at least a 12% performance improvement, significantly enhancing both computational efficiency and portability.