PageWeaver: KV-Guided Query Unions for Sparse Attention

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the low GPU utilization in dynamic sparse attention caused by small query support sets by proposing a page-affinity-based query assembly mechanism. It enables shared KV page loading through bounded search and ID-aware kernels, while an execution-level query union strategy eliminates tensor rearrangement and cross-page partial outputs to optimize non-local reuse. Computational efficiency is further enhanced by integrating FP8 KV caching, Tensor Core tile padding, and dual-CTA kernels. Evaluated on NVIDIA H200 GPUs, the proposed approach achieves a 1.7× speedup over the FlashInfer baseline and improves prefill throughput by 7.88%–14.36%.
📝 Abstract
Dynamic sparse attention limits the KV pages selected by each query, but a small support does not necessarily yield efficient GPU work. Query unions share page loads and populate Tensor Core tiles; their cost depends on which queries are grouped together. We present PageWeaver, an execution design that uses selected-page affinity to assemble query groups while preserving each query's original support and complete output ownership. A bounded GPU search produces query IDs, and an ID-aware two-CTA kernel consumes them without materializing reordered Q tensors or cross-page partial outputs. A direct KV-page union implementation provides a complementary design study of nonlocal reuse and reduction cost. With FP8 KV throughout, the H200 Union8 implementation achieves a 1.70x geometric-mean complete-call speedup over the measured FlashInfer path on six captures. Online regrouping further lowers latency by 3.26-7.66% on five selected 64K-context captures. Whole-model prefill throughput is 7.88-14.36% above the tested native path; the incremental regrouping benefit is smaller, with observed median gains of 0.47-0.73% at 32K/64K and regressions at 8K. A B300 comparison identifies cases where preparation cost and a stronger native kernel remove the advantage. These results separate execution-group reuse from the complete cost of exploiting it online.
Problem

Research questions and friction points this paper is trying to address.

Sparse Attention
Query Unions
GPU Efficiency
KV Pages
Tensor Core Utilization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Attention
Query Unions
KV-Guided
Tensor Core
Online Regrouping
🔎 Similar Papers
No similar papers found.