Sparse inverse Cholesky factorization of dense kernel matrices by greedy conditional selection

📅 2023-07-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the challenge of balancing computational efficiency and approximation accuracy in sparse approximate inverse Cholesky decomposition of high-dimensional kernel matrices. We propose a dynamic sparsity pattern construction method based on greedy conditional mutual information maximization—marking the first integration of mutual information maximization into sparse structure learning, enabling data-adaptive pivot selection. Unlike conventional geometric neighborhood constraints, our approach jointly leverages partial Cholesky updates and KL-divergence minimization to support efficient multi-objective aggregation during decomposition. Theoretical analysis shows the time complexity reduces to *O(Nk²)*, substantially improving upon the *O(N³)* cost of standard Cholesky decomposition. Experiments on Gaussian process regression and image classification demonstrate superior preprocessing accuracy and faster conjugate gradient convergence compared to k-nearest-neighbor and other baseline methods. Our framework establishes a scalable, high-fidelity sparse approximation paradigm for large-scale kernel learning.
📝 Abstract
Dense kernel matrices resulting from pairwise evaluations of a kernel function arise naturally in machine learning and statistics. Previous work in constructing sparse approximate inverse Cholesky factors of such matrices by minimizing Kullback-Leibler divergence recovers the Vecchia approximation for Gaussian processes. These methods rely only on the geometry of the evaluation points to construct the sparsity pattern. In this work, we instead construct the sparsity pattern by leveraging a greedy selection algorithm that maximizes mutual information with target points, conditional on all points previously selected. For selecting $k$ points out of $N$, the naive time complexity is $mathcal{O}(N k^4)$, but by maintaining a partial Cholesky factor we reduce this to $mathcal{O}(N k^2)$. Furthermore, for multiple ($m$) targets we achieve a time complexity of $mathcal{O}(N k^2 + N m^2 + m^3)$, which is maintained in the setting of aggregated Cholesky factorization where a selected point need not condition every target. We apply the selection algorithm to image classification and recovery of sparse Cholesky factors. By minimizing Kullback-Leibler divergence, we apply the algorithm to Cholesky factorization, Gaussian process regression, and preconditioning with the conjugate gradient, improving over $k$-nearest neighbors selection.
Problem

Research questions and friction points this paper is trying to address.

Sparse inverse Cholesky factorization for dense kernel matrices
Greedy selection algorithm maximizing mutual information
Efficient time complexity for multiple target points
Innovation

Methods, ideas, or system contributions that make the work stand out.

Greedy selection maximizes mutual information conditionally
Reduces time complexity via partial Cholesky maintenance
Applies to Cholesky factorization and Gaussian processes
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Carnegie Mellon University | Washington University in St. Louis | University of Wisconsin–Madison | Caltech | Georgia Institute of Technology
S
Stephen Huan
Computer Science Department, Carnegie Mellon University
J
J. Guinness
Department of Statistics and Data Science, Washington University in St. Louis
M
M. Katzfuss
Department of Statistics, University of Wisconsin–Madison
H
H. Owhadi
Computing and Mathematical Sciences, Caltech, Pasadena, CA
F
Florian Schafer
Georgia Institute of Technology, S1317 CODA, 756 W Peachtree St Atlanta, GA 30332