PageWeaver: KV-Guided Query Unions for Sparse Attention
This study addresses the low GPU utilization in dynamic sparse attention caused by small query support sets by proposing a page-affinity-based query assembly mechanism. It enables shared KV page loading through bounded search and ID-aware kernels, while an execution-level query union strategy eliminates tensor rearrangement and cross-page partial outputs to optimize non-local reuse. Computational efficiency is further enhanced by integrating FP8 KV caching, Tensor Core tile padding, and dual-CTA kernels. Evaluated on NVIDIA H200 GPUs, the proposed approach achieves a 1.7× speedup over the FlashInfer baseline and improves prefill throughput by 7.88%–14.36%.