🤖 AI Summary
This study addresses the redundant KV loading and computational overhead in sparse attention caused by independent query processing. We propose an optimization framework that jointly processes adjacent queries to reuse shared KV entries. The core innovation, Shift-Compare Set Decomposition (SCSD), transforms irregular set operations into data-parallel primitives, enabling cascaded sharing and tile-aware execution. This is further integrated with a workload-aware sparse attention mechanism and pipeline optimization strategies. Experimental results demonstrate that our approach reduces kernel latency by 55.1% and KV data throughput by 55.9%, while decreasing time-to-first-token by 36.8%, all with negligible degradation in model accuracy.
📝 Abstract
Sparse attention reduces the cost of long-context attention, but existing kernels typically process queries independently, repeatedly loading and dequantizing KV entries shared across queries. We observe substantial overlap in the KV entries selected by neighboring queries, creating opportunities for cross-query reuse. We present QUILT, a workload-aware sparse-attention execution mechanism that jointly processes neighboring queries and reuses shared KV entries to reduce redundant memory traffic and computation. QUILT introduces Shift-and-Compare Set Decomposition (SCSD), which transforms irregular set operations into regular data-parallel primitives suitable for modern accelerators, and pipelines SCSD with attention computation to hide its overhead. Cascaded sharing captures reuse hierarchically at multiple granularities. A tile-aware execution strategy balances sharing granularity with hardware tile utilization and selectively removes low-importance query-specific tails to eliminate underutilized tiles. We evaluate QUILT on LongBench using GLM-5.3 and DeepSeek-3.2 under both tensor and sequence parallelism. Compared with the state-of-the-art sparse-attention kernel, QUILT reduces average kernel latency by up to 55.1% and processed KV data by up to 55.9%, while reducing time-to-first-token (TTFT) latency by up to 36.8% with negligible accuracy degradation.