🤖 AI Summary
This work addresses the high memory access overhead and KV cache storage costs in the Transformer decoding phase by leveraging the Denotational Normal Form (DNF) and the Mathematics of Arrays (MoA). Starting from DNF, it algebraically eliminates the Kᵀ buffer, achieving theoretically optimal DRAM traffic. By applying MoA-based concatenation, the per-step KV cache append complexity is reduced to O(dₖ + dᵥ), matching the theoretical lower bound. The study also rigorously derives grouped-query attention (GQA) and multi-query attention (MQA) variants, proving they reduce KV traffic by a factor of h_q/h_kv. The proposed GPU kernel achieves zero numerical error with DRAM traffic reduced to (dₖ + n dₖ + n dᵥ + dᵥ)×4 bytes (error ≤ 2×10⁻⁷), and all optimizations are validated against PyTorch benchmarks.
📝 Abstract
We derive four memory-optimal inference artifacts for transformer attention using the Mathematics of Arrays (MoA), each following directly from the forward-pass Denotational Normal Form (DNF) of with the query-row index fixed to the current decode step. The artifacts are: (1)~a single-query decode DNF in which the $ψ$-reduction eliminates the $K^\top$ buffer algebraically, achieving $(d_k + nd_k+ nd_v+ d_v)\times4\,{B}$ Dynamic Random Access Memory (DRAM) traffic result numerically verified to $\|{err}\|_\leq2\times10^{-7}$; (2)~a C/OpenACC Graphics Processing Unit (GPU) kernel with Operational Normal Form (ONF) stride arithmetic and hardware-coalesced memory access, verified to $\|\mathrm{err}\|_\infty=0$ (exact IEEE-754 floating-point arithmetic); (3)~a multi-step KV-cache with $O(d_k+d_v)$ per-step append via MoA concatenation $\#$; and (4)~Grouped-Query Attention (GQA) and Multi-Query Attention (MQA) derived via $ψ$-selection, achieving a proven $\frac {h_q} { h_{kv} }$ reduction in KV traffic. All programs are verified against PyTorch scaled_dot_product_attention.