Fast Evaluation of Truncated Neumann Series by Low-Product Radix Kernels

📅 2026-02-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work proposes a fast recursive framework for efficiently computing the truncated Neumann series $S_k(A) = I + A + \cdots + A^{k-1}$ by leveraging high-radix kernel functions $T_m(B) = I + B + \cdots + B^{m-1}$, substantially reducing the number of required matrix multiplications. The key contributions include the first exact radix-9 rational-coefficient three-product kernel and a residual-based recursive structure that enables stable application of approximate kernels. Furthermore, a radix-15 kernel is designed, achieving the best-known asymptotic complexity to date. Theoretical analysis shows that the radix-9 and radix-15 kernels reduce the multiplication count to $1.58\log_2 k$ and $1.54\log_2 k$, respectively, yielding approximately 21% speedup over conventional repeated squaring. Experimental results corroborate the predicted acceleration.

Technology Category

Machine Learning: Kernel MethodsKnowledge Representation and Reasoning: Computational Complexity of ReasoningSearch and Optimization: Non-convex Optimization

Application Category

Graph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsSecurity and Privacy: Large-scale security measurementsSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search engines
📝 Abstract
Truncated Neumann series $S_k(A)=I+A+\cdots+A^{k-1}$ are used in approximate matrix inversion and polynomial preconditioning. In dense settings, matrix-matrix products dominate the cost of evaluating $S_k$. Naive evaluation needs $k-1$ products, while splitting methods reduce this to $O(\log k)$. Repeated squaring, for example, uses $2\log_2 k$ products, so further gains require higher-radix kernels that extend the series by $m$ terms per update. Beyond the known radix-5 kernel, explicit higher-radix constructions were not available, and the existence of exact rational kernels was unclear. We construct radix kernels for $T_m(B)=I+B+\cdots+B^{m-1}$ and use them to build faster series algorithms. For radix 9, we derive an exact 3-product kernel with rational coefficients, which is the first exact construction beyond radix 5. This kernel yields $5\log_9 k=1.58\log_2 k$ products, a 21% reduction from repeated squaring. For radix 15, numerical optimization yields a 4-product kernel that matches the target through degree 14 but has nonzero spillover (extra terms) at degrees $\ge 15$. Because spillover breaks the standard telescoping update, we introduce a residual-based radix-kernel framework that accommodates approximate kernels and retains coefficient $(\mu_m+2)/\log_2 m$. Within this framework, radix 15 attains $6/\log_2 15\approx 1.54$, the best known asymptotic rate. Numerical experiments support the predicted product-count savings and associated runtime trends.
Innovation

Methods, ideas, or system contributions that make the work stand out.

high-radix kernels
truncated Neumann series
matrix polynomial evaluation
approximate matrix inversion
residual-based framework